Elon Musk's xAI has announced a roadmap to 1.44 million GPU capacity following recent deployments of 660,000 NVIDIA Blackwell GPUs.
Elon Musk's xAI plans to deploy 660,000 Nvidia Blackwell GPUs at its Colossus supercomputer by year-end, raising the total fleet to 1.44 million accelerators, according to a Sept. 25 post on X. The company aims to bring 220,000 GB300 GPUs online by end of September, another 220,000 in November, and a third batch of 220,000 by late December, contingent on hardware deliveries and deployment work.
The expansion would give xAI 1.1 million GB300 GPUs across its facilities. Colossus 1 currently operates 230,000 GPUs—comprising 150,000 H100s, 50,000 H200s, and 30,000 GB200s. Colossus 2 holds 550,000 GPUs, including 110,000 GB200s and 440,000 GB300s. The new additions would bring the GB300 count from 440,000 to 1.1 million, completing a fleet of 150,000 H100 GPUs, 50,000 H200 GPUs, 140,000 GB200 GPUs, and 1.1 million GB300 GPUs.
Nvidia's GB300 NVL72 platform uses Blackwell architecture for large-scale AI workloads. Colossus 2's single-generation design provides a more consistent software and networking environment than Colossus 1, which mixes Hopper and Blackwell systems.
xAI uses uniform accelerators to train Grok, its AI system. Different GPU generations force engineers to tune workloads around the slowest hardware, reducing cluster efficiency and complicating scheduling. Colossus 2's Blackwell-only design creates a uniform training pool, while xAI has reserved Colossus 1's mixed fleet for other workloads, including inference capacity rented to Anthropic. Training and inference place different demands on data centers: training moves large data sets across many accelerators simultaneously, while inference serves completed models and can better utilize varied hardware depending on workload size and response targets.
Power availability has become the central constraint. xAI installed gas turbines near Colossus to provide extra capacity, though residents challenged the project in a lawsuit over permitting. xAI committed to removing the equipment over one year as a 1.2-gigawatt power plant comes online. That plant would supply dedicated electricity for the Colossus expansion. While a 1.2-gigawatt supply cannot sustain every future expansion alone, it would support major growth from the current fleet. Data center operators must secure land, grid connections, generation capacity, cooling systems, and high-speed networking before translating purchased GPUs into usable compute.
Musk has described the 1 million-GPU milestone as an early stage. He has stated the company plans to increase data center capacity sevenfold by 2027 and pursue 50 million H100-equivalent GPUs by 2030. Those figures include future systems that may employ newer accelerators; the H100-equivalent metric compares computing capacity across generations rather than physical unit count. Musk has proposed an orbital data center network with 1 million satellites, though Nvidia CEO Jensen Huang characterized that concept as aspirational rather than an operating project, leaving ground-based facilities as xAI's practical near-term path. Broadcom said in 2024 that three hyperscale customers aimed to reach 1 million GPUs by 2027 without identifying them. xAI's planned 1.44 million-GPU fleet would position it among the largest AI infrastructure operators, contingent on meeting delivery, power, and deployment targets.