Amazon Web Services has reportedly instructed its engineers to reduce their internal use of EC2 instances. This directive, issued in May, aims to ensure sufficient CPU capacity is available to meet customer demand. Engineers who previously accessed instances within hours are now experiencing delays of several days.
The increased demand for CPUs in data centers is attributed to the rise of agentic AI workloads. Unlike traditional AI inference, which is heavily GPU-accelerated, agentic AI involves more complex tasks like tool calls and orchestration that rely significantly on CPU processing. This shift has elevated the importance of CPUs in AI infrastructure.
Historically, AWS engineers could easily provision EC2 instances for development. The current scarcity indicates a significant change in resource availability within Amazon, with one engineer noting unprecedented wait times. This internal pressure highlights the broader industry challenge of managing CPU resources in the era of advanced AI.
The high demand for CPUs is a widespread issue, with major chip manufacturers like Intel and AMD reporting that companies are eager to acquire any available processors. AMD recently launched its Zen 6 'Venice' CPUs for data centers, marking a strategic shift to prioritize the data center market over client markets for new architectures.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
Amazon Web Services is limiting its engineers' access to EC2 instances to conserve CPU capacity for paying customers. This internal crackdown is a direct result of surging CPU demand driven by complex agentic AI workloads, which require more CPU resources than traditional GPU-centric AI inference.