China's AI sector is starting to calculate the economics of compute, as model call prices keep falling while compute investment keeps surging.
*Originally published by Qifan Business and republished by TMTPost with authorization.*
On October 8, the *Financial Times* turned its attention to Ulanqab in Inner Mongolia.
The northern city, several hundred kilometers from Beijing, is becoming one of China's most densely built data center hubs. Local authorities have signed 89 data center projects, and Alibaba, Huawei, Kuaishou, and third-party data center operators have all set up facilities there. The report also noted one detail: a disused local school is being converted into a data center.
Ulanqab attracts internet companies not because of its population or software talent, but because of relatively cheap electricity, available land, and a climate suited to cooling servers.
In July, *China News Weekly* visited the area. On land that was previously dominated by traditional industries, data centers from Alibaba and Sinnet (Zhonglian Data) have begun to appear. The local government has proposed building a "Token Capital," an effort to turn its energy resources into a new industrial advantage.
*Image source: Zhonglian Data, Ulanqab data station*
The construction boom in Ulanqab coincides with a wave of price cuts in the AI industry.
In May, Reuters reported that DeepSeek would permanently cut API prices for its V4-Pro model by 75%. Prices across its billing tiers fell from between RMB 0.1 and RMB 24 per million tokens to between RMB 0.025 and RMB 6.
*Image source: Reuters report on DeepSeek's permanent model price cut*
By October, the *Financial Times*, citing the Token price index compiled by Silicon Data, reported that the composite token price had fallen more than 52% from its May peak. The decline reflected both price adjustments by model providers and companies shifting some tasks to cheaper models. Not all models saw cuts of the same magnitude.
Model services are getting cheaper, yet technology giants continue to buy servers and build data centers at scale.
On August 20, Alibaba released its earnings for the quarter ended June 30. Quarterly capital expenditure reached RMB 67.678 billion, up 75% year on year. Over the same period, revenue from AI cloud and computing services reached RMB 48.437 billion, up 45% year on year.
Taken together, these two figures show a change underway in the AI industry. Model companies keep lowering service prices, while cloud providers must invest heavily up front in computing infrastructure that will take years to pay back.
For the past two years, industry discussion has focused mainly on model capability and GPU counts. The company that trains the strongest model or owns the largest compute cluster has often commanded the most market attention.
As models move into real-world use, a more conventional set of business rules is starting to apply. Servers depreciate, data centers need continuous power, and computing resources must find paying customers. Technological progress can lower costs, but it does not guarantee that every infrastructure investment will earn a satisfactory return.
The competition in Chinese AI is shifting from securing compute to operating it.
## 01 Why Are AI Giants Moving Closer to Power Plants?
Ulanqab did not come onto the internet companies' radar only after large models appeared.
Long before the rise of generative AI, Alibaba, Huawei, and others were already building data centers in regions with favorable energy conditions. Cloud computing changed the commercial nature of servers. In the past, internet companies bought computing equipment mainly to support their own businesses, such as search, e-commerce, and video. With the rise of cloud services, computing, storage, and networking gradually became standardized services that could be sold to outside customers.
Servers became more than a technology investment. They became operating assets that must generate revenue continuously. For cloud providers, beyond equipment purchase prices, the costs of data center construction, power supply, network connectivity, and operations all determine whether the business can make money.
Large models pushed computing demand to a new scale.
Training a model requires mobilizing large numbers of GPUs within a relatively concentrated period, and once a model is deployed, it must keep responding to user requests. AI is now expanding into programming, office work, and enterprise software. Businesses that once needed only a small number of servers are now adding model calls, data processing, and computing tasks.
One of the most obvious costs of a data center is electricity.
Consider a hypothetical data center with an average IT load of 10 MW, running continuously throughout the year, with a PUE of 1.2. If the electricity price differs by RMB 0.40 per kilowatt-hour, the annual difference in electricity costs would be roughly RMB 42.05 million. Over five years, the cumulative difference could exceed RMB 200 million.
This is a hypothetical calculation meant to illustrate cost sensitivity and does not correspond to any specific data center in Ulanqab. Actual electricity costs also depend on equipment load rates, power contracts, cooling systems, and operating methods. Still, the impact of electricity prices on large computing facilities is clear.
If a cloud provider operates several large data centers at once, differences in power procurement terms alone can create a considerable cost advantage. Issues that once seemed to belong to local investment promotion and energy supply are now part of technology giants' investment decisions.
This is where Ulanqab's advantages come into play. The area has relatively abundant electricity, land suitable for large-scale facilities, and a northern climate that can reduce cooling costs. The *Financial Times* reported that some data centers in Inner Mongolia can lower energy costs by using the local power system and direct connections to power plants.
These conditions can help companies reduce long-term operating expenses, but servers cannot simply be placed anywhere that has power.
Training models requires high-speed interconnected compute clusters, while inference services must account for network latency. For products that need fast responses, being close to major users and network hubs still matters. Moving computing facilities to regions with lower energy costs may save on electricity, but it also creates new requirements for network transmission, resource scheduling, and system maintenance.
Large cloud providers must balance these factors. Some training, batch processing, and storage tasks suit the energy-cost advantages of remote sites, while latency-sensitive services need computing resources deployed according to user distribution. Where a data center is built ultimately depends on what the company wants it to do.
Location is only one part of the investment decision.
Whether a company buys a given batch of computing equipment itself or rents it from a cloud provider also leads to different financial outcomes.
A cloud provider such as Alibaba can build data centers, purchase GPUs, and sell computing capacity to external customers. Third-party data center operators provide facilities, power distribution, and operations services, and charge fees under long-term contracts. Model companies can buy cloud services based on actual usage, or they can build their own compute clusters.
Each approach secures computing power, but the risks differ.
Building in-house requires a company to commit capital upfront, and the company bears the risk of whether it can make full use of the equipment. Buying cloud services reduces initial capital expenditure and makes it easier to adjust resource levels, but long-term costs may be higher.
For companies with stable computing demand and sufficient scale, building in-house may be more economical. For AI companies whose businesses are still changing rapidly, leasing can sometimes help them avoid locking in technology and assets too early.
There is also a difference that is easy to overlook.
Buildings, power systems, and network infrastructure can last a very long time, but GPUs may lose their cost competitiveness long before reaching the end of their physical life.
Newer generations of chips usually deliver higher computing performance, and model architectures and inference techniques keep evolving. A GPU a company buys today may still run years from now, but if new equipment can complete the same computing tasks at lower cost, older equipment faces fiercer price competition.
Cheap electricity can improve costs, but it cannot eliminate this risk.
Whether a data center ultimately makes money depends not only on equipment purchase and operating costs, but also on whether it can sustain high utilization and continue selling computing resources at reasonable prices.
In the past, technology companies worried about not having enough chips to train models.
Now, how chips are used after purchase, who bears depreciation, and how much investment can be recovered within a few years are becoming questions that operators must answer.
## 02 DeepSeek Cut Prices by 75%. Why Is Computing Still Not Cheap Enough?
In May, DeepSeek's price adjustment for the V4-Pro model drew attention across the market.
On May 23, Reuters reported that DeepSeek had made permanent the 75% price cut previously applied to the V4-Pro model. According to the information published at the time, API prices across the different billing tiers fell to one-quarter of their former levels.
Earlier, DeepSeek had linked the model's relatively high pricing to constraints on the supply of high-end computing resources, and it indicated that service prices could fall further as relevant hardware supply improved.
But a 75% cut in a model's listed price does not mean the computing cost of completing the same task has also fallen by 75%.
The API price a model company publishes is what customers pay for the service. The actual computing cost a company bears depends on factors including chip performance, model architecture, inference optimization, equipment utilization, and resource scheduling efficiency.
Improvements in cost can create room for price cuts, and market competition can also push companies to lower their fees voluntarily. For customers, both scenarios appear as falling token prices. For model companies, however, the implications for profit may be entirely different.
The Silicon Data index cited by the *Financial Times* on October 5 offers another perspective.
Looking at composite prices, token prices are down more than 52% from their May peak. However, some high-end models have kept relatively stable quotes, and the median price has not fallen by a similar magnitude.
*Image source: Token price observation infographic, Financial Times*
This gap suggests that the way enterprises purchase AI services is changing.
In the past, users tended to treat model capability rankings as the main basis for choosing a product. As model supply grows, more tasks can be handed to cheaper models whose capabilities are good enough. Simple information classification, document organization, and routine customer service do not necessarily require the most expensive products. Complex reasoning, specialized analysis, and high-reliability tasks may continue to rely on high-performance models.
What enterprises care about is how much it costs to complete a task and whether the results meet business requirements.
For model providers, leading capability still has value, but it no longer means customers are willing to pay a premium for every task. Once an ordinary model can meet most of a customer's everyday needs, premium models must prove that their additional performance justifies higher prices.
This shift will gradually affect how profits are distributed across the model industry.
If a model company uses technical optimization to reduce inference costs by 30%, it can choose to keep the efficiency gains as profit, or it can pass part of that room to customers and use lower prices to compete for market share.
When rival services become sufficiently similar, price cuts can become the more effective strategy. Efficiency gains from model companies' R&D investments will, in part, ultimately be passed on to customers through market competition.
For companies using AI, this is good news.
After call prices fall, businesses that were once too expensive to justify AI become economically viable. Some companies first tried models for customer service, copywriting, or office tasks, and later gradually extended AI into software development, data processing, and internal business workflows.
AI agents have also changed the nature of computing demand.
Traditional chat products usually revolve around a single question and answer. Agents, by contrast, try to complete an entire piece of work. An agent may first retrieve information, then call tools, analyze the returned data, and run the task again if the result does not meet requirements. The user submits one request, but the backend may involve multiple rounds of model calls along with additional CPU, storage, and networking operations.
The price of each individual model call keeps falling, but the computing resources needed to complete an entire task do not necessarily decline at the same rate.
Technical optimization allows the same task to be completed with fewer resources.
Lower prices have also attracted more users, and tasks that never previously entered AI systems have begun to consume computing power.
Falling AI service prices and the industry's continued expansion of computing infrastructure are not contradictory.
The harder question is whether this new demand can keep already-deployed equipment generating enough revenue.
When cloud providers buy GPUs, they must forecast computing demand several years ahead. Once equipment is installed in data centers, depreciation, power, and maintenance costs keep accruing, while model prices and computing efficiency can shift significantly every few months.
If a new generation of models can complete the same tasks with fewer computing resources, older equipment will need more orders to maintain utilization. Even if total computing demand keeps growing, the equipment held by any individual cloud provider may not always hold a cost advantage.
Conversely, if AI applications expand quickly enough and more businesses hand complex tasks to models, existing facilities could achieve higher utilization, and equipment installed in advance would more easily recover its costs.
Both outcomes are possible. The difference lies in the pace of technological progress, market demand, and changes in service prices.
Model companies can adjust API prices in response to competition, but cloud providers cannot change capital expenditures that have already been committed at the same speed. The GPUs bought today must be paid back through operations over the coming years.
This mismatch in timing is becoming one of the most difficult problems in AI infrastructure investment.
In the past, companies were willing to pay higher prices for high-performance computing, and cloud providers could use that to estimate the potential revenue of their equipment. As the ways models are used keep changing, cloud providers must simultaneously forecast customer scale, workload structure, model efficiency, and market prices.
A major change in any one of these variables could affect the return on investment.
Model prices can be adjusted quickly, but servers that have already been purchased cannot be returned at will.
The most expensive competition in the AI industry may not take place on model price lists, but after those prices have already fallen.
## 03 Alibaba's Earnings Reveal Another Side of the AI Ledger
On August 20, Alibaba Group released its results for the quarter ended June 30, 2026.
In the quarter, group capital expenditure reached RMB 67.678 billion, up 75% from the same period last year. In its earnings release, Alibaba explained that the increase was related to investment in AI infrastructure, and that procurement cycles, possible CPU demand from AI agents, and rising prices of chip components had also affected the scale of its investment.
Few internet businesses have ever needed to absorb such large amounts of capital in such a short period as today's AI infrastructure does.
However, Alibaba's earnings contain two sets of figures that are worth examining side by side.
In the quarter ended June, revenue from Alibaba's AI cloud and computing services reached RMB 48.437 billion, up 45% year on year, and adjusted EBITA was RMB 5.628 billion, up 133% year on year.
*Image source: Alibaba quarterly results announcement*
Over the same period, revenue from the AI Lab and applications segment was RMB 3.338 billion, with an adjusted EBITA loss of RMB 13.861 billion. Alibaba attributed the widened loss mainly to increased investment in AI capability development and higher inference costs for Qwen app-related services.
These two businesses are at different stages of development and bear different R&D expenses, internal resource costs, and product investments. The gap in adjusted EBITA should not be read as evidence that cloud computing has already recouped its investment while model applications are doomed to struggle with profitability.
Instead, the figures reflect the different operating conditions at different points in the AI industry chain.
For cloud providers, growing computing demand from model companies and enterprise customers translates directly into resource rental revenue. Customers buy computing capacity, cloud providers collect fees under agreed terms, and profits depend on equipment utilization and operating efficiency.
Model companies face a different calculation.
Training models requires purchasing or renting computing resources, and inference services also consume compute continuously. Companies must also bear R&D, product, and marketing investment. Model businesses can only gradually cover these costs when API calls, subscription revenue, or enterprise solutions generate enough paying demand.
The same spending on computing resources can appear as revenue in a cloud provider's financial statements and as a cost in a model company's.
If a model company chooses to build its own cluster, cloud service fees that were previously paid monthly turn into upfront capital expenditure, along with ongoing depreciation and operating costs.
Third-party data center operators play a somewhat different role. They earn fees through facility leasing, power distribution, and operations services. They do not necessarily own the GPUs that customers use, nor do they necessarily participate directly in commercializing model services.
Chip makers, data center operators, cloud providers, and model companies all take part in the expansion of AI infrastructure, but they do not necessarily earn returns at the same time.
Chip makers earn sales revenue when they sell equipment. For cloud providers, buying equipment is only the beginning of the investment, and costs must be recovered through future computing service revenue. Model companies must keep searching for customers who are genuinely willing to pay for AI capabilities.
Continued growth in AI investment does not mean every company in the industry chain will earn higher profits at the same time.
Alibaba's capital expenditure also highlights a problem that operating profit can easily obscure.
Even if the adjusted EBITA of the cloud business improves substantially, the group still needs to keep investing heavily in infrastructure. Operating profit reflects performance under a specific accounting framework and does not mean that new equipment investments have already been recovered.
For cloud providers, whether a computing asset is worth buying depends on how much revenue it can generate over its full life cycle, and how much cost it will consume through electricity, maintenance, depreciation, and equipment replacement.
This differs from some business models that internet companies have long been familiar with.
Once search, social, and e-commerce platforms reach a certain scale of users, they can grow revenue through advertising, transaction commissions, and value-added services. Although these businesses also depend on technology infrastructure, digital distribution often has strong economies of scale.
AI model services, by contrast, require computation to be carried out continuously. Each time a user calls a model, the backend consumes computing resources. Improved inference efficiency can lower unit costs, but it does not eliminate computing investment altogether.
Whether a company makes money depends on whether customers are willing to pay enough to cover the cost of completing tasks, and whether larger usage volumes can raise equipment utilization.
This means technology companies' competitive advantage no longer depends solely on model capability.
A company may have a model that leads in performance yet struggle to scale commercially because its service costs are too high. Another may not have the strongest model on the market, but still earn steady revenue thanks to lower computing costs, more stable services, and a mature customer base.
Cloud providers also face competition.
If several vendors can offer similar computing resources, customers will compare prices, stability, development tools, and migration costs. Scale helps companies meet demand, but it does not necessarily give them stronger pricing power.
Model companies face another test. When low-cost models can handle large volumes of routine tasks, premium models must prove through real-world results that their higher fees are justified. Improving model performance remains important, but customers ultimately buy the ability to get work done, not technical metrics for their own sake.
Whether this round of AI investment earns satisfactory returns will ultimately depend on what happens in real business.
A company may keep buying AI services because software development cycles have shortened, customer service has become more efficient, or certain workflows require less manual processing. If these gains are large enough to cover usage costs, customers have a reason to expand their spending. If a model is powerful but fails to deliver measurable operational improvements, companies will find it hard to keep increasing their budgets indefinitely.
The data center construction in Ulanqab, DeepSeek's price adjustments, and Alibaba's capital expenditure appear to belong to different parts of the industry. In reality, they point to the same question.
Low electricity prices can reduce operating expenses, but they cannot guarantee that a data center will have enough customers. Price cuts for models can broaden usage, but they may also compress the pricing room for some services. Growth in cloud revenue can improve operating performance, yet newly added computing assets still take years to recover their investment.
China's AI industry remains in a phase of investment expansion, and computing demand still has room to keep growing. But growth in industry demand does not guarantee that every company's servers will maintain a cost advantage over the long term, nor that every investment will earn a sufficiently high return.
Over the past two years, AI companies needed to prove they could train stronger models. As computing infrastructure investment grows heavier, asset management and capital returns are becoming part of the competition alongside technical capability.
Model prices can fall by three-quarters within a few months, while a data center investment may take years to recover.
When technology iterates in months but capital is recovered over years, what Chinese AI companies may most need to learn is not only how to train better models, but how to run a heavy-asset business.