Chinese compute providers are shifting focus from raw GPU stacking to optimizing Token value output per unit of infrastructure as resource constraints tighten.
In 2026, a near-frenetic infrastructure expansion is underway across major data center hubs in northwest China. In June, DeepSeek placed a massive bet on a 1-gigawatt intelligent computing base in Ulanqab, Inner Mongolia. By July, Sugon 8000 announced its first fully domestic 100,000-card supercluster officially connected to the National Supercomputing Internet. In August, Alibaba Cloud declared it had compressed its AI Data Center (AIDC) delivery cycle to just 100 days, aggressively expanding along the northwest’s green energy corridor. Notably, unlike previous vendor efforts focused on tallying hardware units or showcasing theoretical peak compute, the industry has pivoted toward systems engineering. The true benchmark for computing infrastructure appears to have shifted from “how many cards you own” to “whether 100,000 cards can stably, continuously, and without bottlenecks collaborate to execute a real-world large model task.”
This evolution raises a critical question: As computing infrastructure transitions from simple hardware stacking into a grueling systems engineering test, what is the fundamental driver behind this shift? Why does the industry collectively face anxiety over “effective output” even as card counts soar? The answer lies in downstream demand exploding beyond supply. Over recent years, AI deployment has evolved from Chat AI to continuous model-calling Agents, driving a surge in model invocations. According to the National Data Administration, China’s daily average token call volume exceeded 140 trillion in March 2026—a more than 1,000-fold increase from the 100 billion recorded in early 2024. This demand surge is reshaping the global landscape of model usage. Data from OpenRouter shows that between February 9 and 15, 2026, Chinese models accounted for 4.12 trillion token calls, surpassing the 2.94 trillion calls for U.S. models during the same period.
The mechanics are straightforward. Chat AI typically operates on a one-question-one-answer basis, whereas Agents autonomously plan tasks, invoke tools, parse results, and iterate through execution and verification. A single task often requires multiple consecutive runs—sometimes dozens—driving token consumption far higher than standard conversations. Simply put, the tasks fueling this surge are becoming significantly heavier. A study on SWE-bench Verified using eight frontier models and the OpenHands Agent revealed that agentic coding consumes approximately 3,500 times more tokens than single-turn code reasoning and 1,200 times more than code chat. Input tokens heavily outweigh outputs, yielding an average input-to-output ratio of 154:1.
As these high-consumption tasks scale from experimental use to mass application, platform-level token structures are shifting. The OpenRouter and a16z “2025 AI Usage Report,” based on anonymized metadata exceeding 100 trillion tokens, shows programming tasks now account for over 50% of platform token usage, up from 11% in early 2025, making it the largest single category. Meanwhile, the most prominent general-purpose Agent framework of March 2026, OpenClaw, contributed only about a quarter of weekly platform consumption. Currently, it is not generic Agents but AI Coding—which has established clear workflows and proven willingness to pay—that reliably converts high-frequency calls into stable token consumption.
Zhipu AI was among the earliest model vendors to concentrate resources on the coding trajectory. Its GLM Coding Plan has surpassed 242,000 global paid developers, with token calls increasing 15-fold over six months. Strikingly, in Q1 2026, GLM API prices were raised cumulatively by approximately 83%, yet call volumes still grew by roughly 400%. Its model API annual recurring revenue (ARR) hit 1.7 billion RMB in March (a 60-fold year-over-year increase) and surged to $1 billion by July. While Anthropic took 15 months to grow from $100 million to $1 billion in ARR, Zhipu achieved it in just five months. Moonshot AI’s growth curve is even steeper. Just 48 hours after launching K3, its user request volume neared the carrying capacity of existing clusters. MiniMax M2.5 broke 3.07 trillion token calls within seven days of launch.
While model vendors witness explosive product demand, cloud providers see the entire token market rapidly expanding. Calculated by MaaS call volume, Volcano Engine’s Doubao large model saw daily token calls rise from 2 trillion at the end of 2024 to 63 trillion by the end of 2025, reaching 120 trillion in March 2026 and 180 trillion by June—a growth of over 1,500 times in two years. IDC data indicates Volcano Engine holds a 9.5% share of China’s public cloud MaaS market by call volume. Internally, its fastest-growing segment is the coding tool Trae, whose daily token consumption scaled from an initial 8 billion to the 300–400 billion range. In Q4 of fiscal 2026, its AI-related revenue reached 8.971 billion RMB, accounting for over 30% of external commercial revenue for the first time. The Bailian platform’s client base grew eightfold year-over-year, with daily token revenue increasing roughly 15-fold over the past five months. Alibaba Cloud’s financial reports repeatedly cite Bailian model calls and AI coding products like Qoder as key growth drivers.
Ultimately, Agents have unlocked a token consumption space far exceeding traditional chat, and AI coding has successfully converted this demand into stable revenue. This has overwhelmed supply, pushing the AI industry into a golden boom phase characterized by “demand explosion → price increases → scale expansion.” Normally, stronger demand yields greater economies of scale and synchronized revenue/profit growth. Yet in the Agent and AI coding era, a paradox has emerged: the hotter the product, the more frequently vendors restrict purchases, raise prices, or suspend new user sign-ups. Moonshot AI exemplifies this. Within 48 hours of releasing K3, compute overload forced it to pause C-end subscription access. If new demand reliably translated into profit, vendors would have no reason to “close shop” during peak demand. Bai Wenxi, Vice Chairman of the China Enterprise Capital Alliance, noted, “The core issue is that K3 passed the demand test but failed the supply and profit test.” Other experts attribute this to service guarantees following compute shortages, reflecting how independent large model companies are shifting from user acquisition to prioritizing revenue and efficiency.
Alibaba Cloud’s Bailian Coding Plan Lite halted new purchases on March 20, 2026, and stopped renewals and upgrades on April 13, phasing out its former 40-RMB low-tier option. Zhipu’s GLM shifted to fixed-time limited sales, with the Max tier releasing only about 20% of its quota daily. Concurrently, Alibaba Cloud, Tencent Cloud, MiniMax, and Xiaomi have introduced or rebranded offerings as Token Plans, replacing per-request billing with granular Credits/Token metering. Previously, Chat AI token consumption was relatively flat and predictable. In the Agent era, particularly for AI coding, this logic breaks down. As noted earlier, certain agentic coding tasks exhibit input-to-output ratios as extreme as 154:1. This rapidly widens the cost gap between users. A casual user might make dozens of calls daily, while a heavy user deeply integrating a coding Agent into their workflow could keep the model running continuously for hours or all day. Individual token and compute costs can thus diverge by tens or even hundreds of times.
This creates a structural problem. Monthly subscriptions essentially use fixed revenue to cover unpredictable compute consumption—a model Agents amplify to the extreme. Stronger products attract more professional developers and high-frequency users. The more thoroughly they utilize the system, the higher the vendor’s marginal compute costs become. Once pre-set capacity is breached, vendors must resort to quotas, limits, benefit splitting, or token-based billing to regain cost control. If token gross margins were exceptionally high, vendors could cross-subsidize heavy users with light ones. In reality, margins are already razor-thin. Sun Yuanhao, CEO of Transwarp, stated, “The root cause of fierce competition in the token factory track is the narrow price spread between token pricing and compute procurement costs, leaving minimal profit space.” Coupled with persistently high hardware procurement costs and multiple rounds of MaaS price wars, the absolute gross margin per token is extremely low. When heavy users trigger long-context retrievals via Agent tasks, these already slim margins vanish instantly, sometimes flipping to net losses. Beyond front-end business model inversion, extreme backend compute waste further inflates the amortized cost per token. Data from the China Mobile Communications Research Institute shows traditional AI center GPU utilization typically remains below 30%. Compute is not efficiently absorbed by the market; vast server farms run idle tasks or stall due to network packet loss and VRAM fragmentation. These idle and wasted cycles are ultimately amortized into every generated token, exacerbating the already squeezed price spread.
This defines 2026’s most paradoxical AI industry scene: booming usage desire and soaring call volumes clash with vendors lamenting shrinking margins and forced price hikes. The underlying cause is that traditional SaaS monthly subscription logic has completely failed against the compute black hole of the Agent era. Demand is real, but given current comprehensive token production costs, no vendor can sustain heavy user demand surges under existing business models. To win this cost defense war, the supply side has launched a systematic overhaul across four pillars: lowering electricity costs, clearing communication pathways, implementing smart division of labor, and capturing off-peak windows. Industrial power in eastern coastal data centers typically ranges from 0.6 to 0.8 RMB per kWh. In contrast, northwest regions like Ulanqab, Inner Mongolia, and Qingyang, Gansu, possess abundant wind and solar green power, bringing comprehensive electricity costs down to 0.25–0.3 RMB per kWh. For AI centers spanning hundreds of megawatts to 1 GW, this slashes energy costs at the source.
However, cheap power does not guarantee cards run at full capacity. In large-scale clusters, communication bottlenecks, network congestion, data packet loss, and node failures force GPUs into idle waits. The larger the scale, the more any minor bottleneck amplifies, resulting in “many cards, but few actually working.” Consequently, architectures like the first fully domestic 100,000-card supercluster and Huawei’s Atlas 950 SuperPoD focus on widening the “road.” Through faster interconnects and tighter node coordination, they transform dispersed servers into a unified computational system, reducing communication latency and ensuring more GPU cycles are dedicated to actual computation. If super-nodes solve “how cards collaborate,” software layers address “what each card should do.” Facing the extreme 154:1 input-output ratio of Agent scenarios, assigning the same GPU to both ingest massive context and generate output leads to severe mismatch. Runtime systems like SenseTime Da Zhuang and Zhongke Jiahe implement “pipeline division of labor.” They assign high-compute, fast-read GPUs to ingest vast historical contexts, while dedicating large-memory, high-bandwidth GPUs to rapidly output code. This specialized “assembly line” operation eliminates VRAM fragmentation and compute bubbles, drastically boosting the cluster’s actual Model Flops Utilization (MFU).
Even with cheap power and fast interconnects, if daytime racks are maxed out while nighttime sees half-empty facilities, machine depreciation remains steep. Compute scheduling has thus evolved from “send work wherever cards exist” to “match the right task, at the right time, to the right location.” Millisecond-response chats stay at eastern nodes near users. Multi-minute or hour-long Agent tasks route to western mega-clusters. Non-real-time jobs leverage nighttime valleys and surplus power. Platforms increasingly use night discounts and elastic token pricing to guide users toward off-peak usage. High-value real-time tasks consume premium resources, while delay-tolerant workloads absorb cheaper, idle compute. This combination of northwest green power, super-node communication breakthroughs, prefill/decode (PD) decoupling, and cross-regional/temporal scheduling transforms scattered, inefficient GPU stacking into a high-throughput, highly utilized “industrial token factory.”
Beneath this grand AI computing expansion campaign lies a highly rational “war to drive down token production costs.” Only when the comprehensive cost per token falls to utility-grade levels will the massive Agent-era demand transition from a heavy-user spree into a sustainable business. As the compute competition’s endpoint narrows from “stacking GPU quantities” to “minimizing the comprehensive output cost per token,” this foundational logic shift will ripple upstream and downstream, restructuring the entire industry chain. Every player—from chip giants and cloud providers to model unicorns, software infra vendors, and enterprise clients—will be drawn into a survival race for “effective compute.” Chipmakers like Huawei, Sugon, Hygon, and Moore Threads are evolving from “selling single cards/theoretical TFLOPS peaks” to “selling software-hardware integrated super-nodes and AI-hardware fusion systems.” As single-card performance gains plateau, network interconnects and ecosystem adaptability will determine survival. Hardware will accelerate toward cabinet-style, pooled super-nodes, deeply binding with upper-layer compute scheduling and heterogeneous compilation frameworks. Cloud providers and AI center operators like Alibaba Cloud, Tencent Cloud, and regional AIDCs are shifting from “selling compute” to “low-cost production and refined token sales.” Alibaba and Tencent compete on extreme delivery efficiency and northwest green power layouts, while simultaneously dismantling monthly plans in favor of Token Plans. Operators like Parallel Computing achieve over 90% utilization through precise scheduling, sharply contrasting with the sub-30% utilization of traditional regional facilities.
Divergence will intensify. SME AI centers unable to join unified scheduling networks or maintain low utilization will face elimination or consolidation waves. Leading clouds will evolve into high-throughput token production networks integrating green power, super networks, and dynamic scheduling. Independent model vendors like DeepSeek, Zhipu AI, and Moonshot AI will gradually sink into becoming builders and operators of full-stack infra. This trend is already visible: DeepSeek plans a 1GW self-built AI base in Ulanqab; Zhipu is deploying a 1GW domestic AI center and directly acquired Zhongke Jiahe, a compilation team originating from the Chinese Academy of Sciences. These moves aim to reduce reliance on external supply and internalize GPU-to-token cost control. Consequently, full-stack vertical integration is becoming the threshold for top Frontier AI enterprises. Only model companies securing cheap power, modular factory construction, and heterogeneous compilation optimization will survive; smaller vendors lacking low-cost token production capabilities will become contract manufacturers or shutter. As model vendors sink deeper into infrastructure, the value of infra and compute software vendors like Zhongke Jiahe, SenseTime Da Zhuang, Jiuzhang Yunji, and iSoftStone is being re-amplified. These firms are transitioning from simple API routing to deep specialization in compilers, Runtimes, heterogeneous mixing, and PD decoupling engines. For instance, SenseTime Da Zhuang lifts mainstream domestic chip MFU by 85%–152% through heterogeneous mixed inference, expanding token output by 2.5x at equivalent costs. iSoftStone and Jiuzhang Yunji reconstruct compute into end-to-end “token factories,” using software engineering to squeeze every drop of performance from each card.
“Software-defined compute” will emerge as a critical value driver. Heterogeneous chip adaptation and distributed compilation capabilities will become extremely scarce assets, likely triggering an industry-wide M&A wave targeting bottom-layer infra teams within the next one to two years. This cost restructuring will inevitably transmit to the downstream of the industry chain. Enterprises and developer clients will be forced to accept granular pay-per-use, per-token, and per-credit billing. Enterprises will intensely monitor token procurement costs, favoring cost-effective domestic tokens or self-hosted private fine-tuned models. Procurement logic will shift accordingly. “Pay-for-performance” and multi-model hybrid scheduling will become standard. Enterprises will establish refined “token cost control systems,” routing simple rule-based tasks to low-cost tokens and reserving high-priced frontier models for highly complex refactoring and debugging. In 2026, the core of China’s compute competition will pivot to the extreme extraction of “effective output” per watt, per card, and per second of depreciation. Companies that engineer unit token costs down to utility-grade levels through systems engineering will truly bridge the profitability chasm and secure their tickets to the Agent era. Players trapped in low-utilization, high-cost GPU-stacking paradigms will ultimately be washed away by the industrial token wave.