Alibaba, ByteDance, and DeepSeek race to build 100-trillion-parameter large language models, escalating China's AI capex arms race
At Alibaba's Cloud Summit on September 22, the company refocused the industry on the "large parameters" trajectory for large models. CEO Wu Yongming disclosed that Alibaba is preparing to train a new generation of AI models with parameter scales between 50 trillion and 100 trillion. In a follow-up statement at the same event, Liu Dayiheng, head of Alibaba ATH's Token Foundry Qwen LLM project, noted that Scaling—the scaling law—is the key path toward artificial superintelligence (ASI), with Qwen 4.5 and Qwen 5 planned to scale toward 50 trillion to 100 trillion parameters.
On the same day, another development rippled through the industry: according to The Information, DeepSeek is currently training a new model with approximately 200 billion parameters. DeepSeek founder Liang Wenfeng also indicated at a recent investor meeting that the company plans to push the next model to 80 trillion parameters.
Two months earlier, in August, the Financial Times of Britain reported, citing sources, that ByteDance was training a large model that could reach up to 100 trillion parameters.
Parameters can be loosely understood as a large model's "brain capacity"—the total volume of experienced weights invoked when the model makes inferences. What does 100 trillion look like? Compare it to DeepSeek-R1, which sparked industry excitement in 2025 with only 671 billion parameters—less than one-fifteenth of 100 trillion.
From Alibaba to DeepSeek to ByteDance, China's large model companies have collectively revived the faith that "bigger is better" in parameters.
Yet the other side of the scale race is this: the larger the parameters, the higher the resources and costs required, yet intelligence does not scale proportionally. Once a model scales up, the data must scale alongside it—and not just quantity, but quality too. When everyone is stacking parameters, the real dividing line becomes computational efficiency: how much intelligence can a single unit of compute buy you.
Alibaba, DeepSeek, and ByteDance are gambling on 100 trillion parameters.
Among published models, the pinnacle of China's lineup is Moonshot's Kimi K3, released in July, with 2.8 trillion total parameters—the world's first open-source model in the 3-trillion range. Alibaba's flagship Qwen 3.8-Max stands at approximately 2.4 trillion, while DeepSeek's current V4-Pro is 1.6 trillion.
Looking overseas, closed-source models show even higher parameter scales. Industry estimates put Anthropic's Mythos 5 at around 80 trillion parameters and Fable 5 at around 50 trillion.
In other words, the highest published parameter models from China still lag the estimated top tier overseas by one to two multiples. And if 100 trillion is the target, each company's current flagship needs to expand roughly five-fold more.
So the question becomes: is bigger really always better?
The logic supporting "bigger is better" points to the Scaling Law—the principle that when model, data, and compute scale up together, model capabilities improve accordingly. Over recent years, from hundred billions to trillions, each parameter leap has brought visible capability jumps. In theory, larger parameters do mean better performance.
This is the confidence behind Liu Dayiheng calling Scaling the key to Alibaba's path to ASI.
But the Scaling Law is not infinitely powerful. It requires preconditions: as parameters grow, data volume and compute must scale proportionally—and crucially, that must be "high-quality data."
At the trillion scale, simply stacking parameters shows sharply diminishing marginal returns. So from a practical standpoint, a so-called 100-trillion-parameter model that cannot be fed an equivalent volume of high-quality text will not see capability improvements scaling up proportionally.
Other variables matter too: data, training methods, architecture, and inference-time algorithms.
For instance, Kimi K3 has 2.8 trillion total parameters, but uses a Mixture of Experts (MoE) architecture in which 896 experts only activate 16 at a time, leaving only about 104 billion parameters actually active. In other words, nominally 2.8 trillion, but far fewer working in practice.
In real use, more parameters doesn't equal "more usable."
On one hand, larger models run slower inference, visibly lengthening user wait times. On the other, inference costs climb, eventually passing through to price—even flagships with only hundreds of billions or a couple trillion parameters already impose substantial user costs.
So setting aside the "is it usable" question and returning to the race itself: parameter scale determines the capability ceiling, and that ceiling determines rank. This is why Alibaba, ByteDance, and DeepSeek, knowing about diminishing returns and soaring costs, are still racing toward 50 trillion to 100 trillion—they are not competing on today's experience, but on the entry ticket for the next generation of models.
The 100-trillion bill: who can afford the cost of large model burnout?
Supermodels can reach higher intelligence, but a bigger foundation means more training and deployment resources. That large models burn money is an unspoken consensus among all major players.
Overseas, Microsoft, Google, Meta, and Amazon have guided 2026 capital expenditures to approximately $720–$745 billion.
China's market is similarly aggressive: Alibaba's Q2 2026 capital expenditure alone was 67.7 billion yuan, up 75% year-over-year, with free cash flow swinging from positive 73.8 billion yuan to negative 46.6 billion yuan. ByteDance's 2026 capex ceiling could reach as high as $70 billion—roughly three times the prior year—and might push another $100 billion in 2027.
Where does the money go? The most direct line is infrastructure.
When training GPT-4, OpenAI needed compute measured in hundreds of billions of tokens, with tens of thousands of GPUs coordinating in high-speed interconnected clusters for weeks or months at a time—and GPT-4 only had 20 trillion parameters.
Pretraining 100 trillion parameters requires GPU counts and durations that scale up by orders of magnitude: tens of thousands of high-end GPUs running continuously for months, with power bills, cluster maintenance, and chip procurement each reaching astronomical figures.
Looking just at data center investment: at the Cloud Summit, Wu Yongming announced that Alibaba Cloud will operate global data centers exceeding 20 GW by 2032. Goldman Sachs previously estimated Alibaba's current online capacity at roughly 3–4 GW. That means adding 16–17 GW over six years, expanding capacity more than fivefold.
Using NVIDIA's Huang Renxiong's disclosure in earnings calls, building 1 GW of compute capacity costs roughly $50–60 billion. By that math, Alibaba's planned 16–17 GW expansion over six years could require total construction investment of $800 billion to $1 trillion.
Next is inference cost. Model training is not the end; it is the beginning. The larger the model, the more compute required for each response. When deployed across millions of daily users, inference-side power and compute expenses often outlast and outspend training.
Then there is a hidden bill often overlooked: data. As models scale, training data must scale too—not just quantity, but quality. Industry consensus is clear: high-quality, finely cleaned and annotated data is now scarcer than raw compute.
Three bills are on the table. Can the players afford them?
Cash flow has already sounded an alarm. Over the past five quarters, Alibaba's free cash flow was negative in four. The latest quarter saw capital expenditure at 67.7 billion yuan. On-hand cash and liquid investments of roughly 474.5 billion yuan provide a cushion, and last month the company raised about 80 billion Hong Kong dollars through a Hong Kong listing placement for AI infrastructure.
But faced with trillion-dollar-scale compute investment, this ammunition is far from enough; sustained external financing seems inevitable.
ByteDance tells a similar story. According to The Information, ByteDance's first-half 2026 net profit declined by single-digit percentages to around $20 billion. It is China's tech giant with the largest AI capital expenditure, with 150 billion yuan in AI-related investment in 2025 and a further increase to 200 billion yuan in 2026.
Facing the gap, ByteDance recently completed a syndicated loan of roughly $30 billion—the largest in tech company history—with some funds expected to flow into ever-expanding AI infrastructure.
An unstoppable arms race: from Flash to 100 trillion parameters.
Less than a month before the Cloud Summit, Alibaba had unveiled Qwen 3.8-Flash-Next: 125 billion total parameters, activating only 6 billion per inference, training cost slashed nearly 90% from the previous generation, with API pricing at as low as one-third of DeepSeek-V4-Flash.
The same day, Zhipu's GLM-5.3-Flash launched on the open-source leaderboard: 320 billion total parameters, activating 18 billion.
Over those two weeks, the industry's rallying cry was "cost revolution": the decisive factor in model competition is not who is biggest but who is thriftiest, fastest, who can compress unit costs lowest in high-frequency use.
A month later, the Cloud Summit put "big" back center stage, and two approaches began to coexist.
From a model standpoint, Flash-class models optimize for "high-frequency, real-time, cost-sensitive" scenarios where users want speed and economy, so active parameters are kept minimal. Large-parameter models, by contrast, optimize for the hardest, most diverse scenarios demanding the toughest reasoning, the longest agent chains, the broadest capability coverage.
What allows these two approaches to coexist and even switch between them is MoE architecture—Mixture of Experts—which divides a model into many "experts," activating only a subset per inference. Its core contribution is complete decoupling of total and active parameters: total parameters can keep ballooning, maintaining the capability ceiling, while per-inference compute stays controlled, not scaling with overall size.
For large model vendors, MoE reshapes the cost structure itself: training cost is manageable, inference charges only for active parameters, effectively letting one foundation bet on both small-and-fast and large-and-strong.
Flash and large parameters ultimately answer the same question: how to extract more intelligence from every dollar spent.
Yet large parameters remain a promise; 100 trillion is still just a slide deck number. China's current flagship peak remains 2.8 trillion. By the time these models ship, the table may already have been reshuffled. Then 100 trillion stops being the finish line and becomes the new starting line.