Huawei shelves global AI chip rollout due to domestic demand outstripping supply; 15,488-chip Atlas clusters achieve 120 EFLOPS using optical interconnect.
Huawei's next-generation Ascend 900-series AI accelerators will be offered only in China, not internationally, as the company struggles to meet domestic demand amid capacity constraints. The upcoming Ascend 960-series neural processing units (NPUs) could rival some of AMD's and Nvidia's existing AI GPUs, though demand for these units outside of China was not guaranteed anyway.
"Since we do not have enough capacity to even satisfy the demand in China, we do not have a plan to expand into the international market in a fully-fledged way," said Eric Xu, rotating chairman of Huawei, on the sidelines of the company's Huawei Connect conference. He added that Huawei supplies limited volumes to "some countries where demand is particularly strong," though he did not elaborate.
Huawei unveiled its latest AI accelerator roadmap this week, revealing major training and inference performance gains for its next-generation Ascend 960, 970, and 980 NPUs over the existing Ascend 910C and Ascend 950-series. The Ascend 960DT and 960PR are set to increase their FP8 training performance to 2 PFLOPS and their FP4 inference performance to 4 PFLOPS and 8 PFLOPS, respectively, in 2027. Their successors, the Ascend 970 and Ascend 980, are projected to reach 3.6 and 7.2 PFLOPS FP8 performance respectively, with FP4 performance reaching 14 PFLOPS and 28 PFLOPS in the coming years.
While the upcoming Ascend NPUs will be considerably faster than their predecessors, particularly for inference, they will remain well behind Nvidia's previous- and current-generation accelerators in raw compute performance. Huawei's 2027 Ascend 960DT is projected to deliver 2 PFLOPS FP8 for training, compared with Nvidia's 4 PFLOPS FP8 for the H200, released in 2023. The Ascend 960PR is expected to offer 8 PFLOPS FP4, which is far behind Nvidia's B300, which delivers 15–20 PFLOPS FP4. Even the Ascend 980, targeted for 2029, is projected to reach 7.2 PFLOPS FP8 and 28 PFLOPS FP4—well below Nvidia's R200, which is on track to deliver 17.5 PFLOPS FP8 and 35–50 PFLOPS FP4.
Such a massive performance differential with leading AI hardware will reinforce Huawei's reliance on massive system-level scaling rather than chip-for-chip performance to compete with Nvidia. However, massive system-level scaling comes with substantial power consumption, which will make Huawei's next-generation Atlas SuperPoDs and SuperClusters considerably less competitive in markets that can access hardware from AMD or Nvidia.
Huawei faces an interesting paradox. On one hand, its integration efforts like near-package optics (NPO) clearly free up capacity on "older" nodes that can be used for other components of AI platforms. On the other hand, SMIC's inability to ramp production on 7nm and 6nm-class nodes limits Huawei's ability to supply its AI hardware, which is why it can barely meet domestic demand.
While Huawei's Atlas SuperPoDs with up to 15,488 Ascend 960 NPUs can deliver up to 30 EFLOPS FP8 and 120 EFLOPS FP4 performance—far exceeding the capabilities of Nvidia's NVL72 clusters—their performance-per-watt is poised to be dramatically lower compared to Nvidia's architectures, which means demand for such hardware outside of China will be limited at best. A global AI hardware push does not make much sense for Huawei right now. What perhaps does make sense is offering cloud access to its hardware to various academic and research customers to popularize its CANN software stack.