Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeHeadlinesReport
Headlines · Report

Amazon Web Services began enforcing internal EC2 CPU restrictions as agentic AI demand from internal projects competes with external customer capacity.

AWS capacity crisis driven by internal AI development; first public sign of hyperscaler saturation and internal/external resource conflict.
Trade pressSlicast · August 8, 2026 · Global · Source: Tom's Hardware
importance 92

Amazon Web Services is cracking down on internal use of EC2 instances among its engineers. In May, the company reportedly met with engineers and told them to reduce CPU waste to ensure AWS has enough CPU capacity to meet customer demand, according to The Information. This message comes as demand for CPUs in the data center has hit a fever pitch, with the traditional eight-to-one or four-to-one ratio of GPUs to CPUs moving closer to parity.

EC2 instances make up a large chunk of the modern internet and are used in both public and private deployments. Traditionally, AWS engineers have been able to spin up their own instances for development, leveraging the relatively low CPU utilization required for web infrastructure to provision more virtual machines. Now, engineers report waiting days to gain access when they previously could obtain access within hours. One engineer told The Information that they had never had to wait this long for an instance, even after several years of working at Amazon.

Amazon deploys several different types of CPUs in EC2 instances, including AMD and Intel options as well as its relatively new Graviton5 chip, Amazon's most powerful CPU to date. The Graviton5 uses an Arm-based architecture along the lines of Nvidia's Vera CPU and Arm's own AGI.

The increased demand for CPU capacity stems from AI agents, a new paradigm in productivity that even companies as large as Amazon are struggling to manage. Last month, a coding agent ran up $1.8 million in token costs at Amazon, surpassing a development budget by 860 percent. Amazon employees have admitted to using AI unnecessarily to pump up internal usage scores, contributing to the strain on capacity.

Much of the AI infrastructure currently in place is designed around inference, a workload accelerated by GPUs. With the traditional four-to-one ratio, the CPU simply served to keep the GPUs fed. However, agentic workloads are considerably more complex. They often involve tool calls that run on the CPU as well as more intricate orchestration of inference on GPUs. This is what has brought CPUs center stage in the agentic era. Intel, AMD, and others have reported what memory and storage companies have also noted over the past several months: demand for CPUs is so high that most companies will take whatever they can obtain. According to reports, Intel has told PC makers to adopt 18A CPUs or risk losing their supply allocation.

The major hardware players are capitalizing on this surge in demand. AMD recently unveiled its portfolio of Zen 6 'Venice' CPUs for the data center, marking the first time AMD has launched a new architecture in the data center before the client market in decades. Nvidia has also pivoted its messaging away from accelerators and toward its new Vera CPU, seeking to claim a position in the expanding market of agentic AI infrastructure.

Although Amazon is wrestling with CPU capacity across both its internal engineers and external customers, The Information reports that shortages are largely confined to spot instances. According to a consultant quoted in the report, contracted capacity has not experienced shortages.

Read the original
Amazon Web Services began enforcing internal… · Slicast