Friday, September 11, 2026
AI 인프라 · 뉴스 & 분석
컴퓨트·클라우드리포트
컴퓨트·클라우드 · 리포트

지푸는 토큰을 티몰을 통해 판매하는 반면, 다른 AI 기업들은 토큰 가격 하락에도 불구하고 공급을 제한하기 시작해 근본적인 컴퓨팅 제약에 대한 의문이 제기되고 있다.

이러한 공급 조정으로의 전환은 GPU 가용성 축소와 추론 비용 압박을 부각시키며, 모델 개발자로 하여금 컴퓨팅 예산을 더 엄격하게 관리하도록 요구하고 있다.
업계 전문지Slicast · September 7, 2026 · 미국 · 출처: PANews
중요도 65

Recently, a reader shared an announcement from Zhipu AI regarding the discontinuation of its long-standing subscription package. The previous “unlimited weekly quota” GLM Coding Plan will cease auto-renewal, with Zhipu compensating legacy users with two months of the new tier free of charge. Just a few months ago, the AI coding market was competing on who offered the lowest prices. Now, the metric has shifted entirely: how many tokens you can actually consume.

On September 2, Zhipu AI officially launched a flagship store on Tmall, listing its GLM Coding Plan. Pricing tiers include the Personal Lite plan at 118 yuan monthly, Pro at 538 yuan, and Max at 1,078 yuan, each tied to distinct credit quotas. Built on GLM-5.3, the service integrates with over 20 mainstream agents, including ZCode, Claude Code, and Codex. What was once a highly technical resource now sits on e-commerce shelves: AI usage quotas can be purchased directly, much like mobile data plans.

Yet the headline isn’t merely that Zhipu AI is selling tokens on Tmall. It’s a counterintuitive industry pivot: tokens are becoming cheaper, yet AI providers are increasingly rationing supply.

Over recent months, a clear shift has emerged across the AI sector. Kimi suspended C-end subscriptions due to compute constraints. On July 19, Moonshot AI announced that demand for its newly launched K3 model far exceeded projections, pushing against existing cluster capacity limits. To prioritize compute for existing paying users, it halted new consumer subscriptions. Similarly, Alibaba Cloud paused new purchases of its Lite Coding Plan on March 20, followed by a halt on renewals and upgrades on April 13. Tencent Cloud adjusted Hunyuan model pricing in March, notably raising the input cost for HY2.0 Instruct from 0.0008 yuan to 0.004505 yuan per thousand tokens—a jump of over 460%.

Why has AI suddenly curbed its earlier generosity? For two decades, internet products have demonstrated that digitization drives marginal costs toward zero. Mobile data exemplifies this: once infrastructure is built, transmitting an additional gigabyte costs almost nothing. As networks scale, prices fall. Many assumed AI would follow the same trajectory—stronger chips, more efficient models, lower inference costs per token, and thus cheaper services.

That premise holds partially true. Zhipu AI disclosed that through continuous optimization of model architecture and inference infrastructure, its unit token inference cost has fallen by 80% year-to-date. The company also confirmed it has achieved large-scale, low-cost inference using hundreds of thousands of domestic chips. The catch? While tokens are cheaper, consumption rates are accelerating even faster. This is the central tension in today’s AI industry.

A token is not ordinary data traffic. Each one represents consumed compute resources: GPUs or AI accelerators, high-bandwidth memory, networking, storage, and the associated power and cooling within data centers. Tokens function as the metered unit for trading AI compute. When a model simply chats, a single request may use hundreds or thousands of tokens. But when tasked with complex workflows, consumption scales dramatically.

Zhipu AI’s latest semi-annual report captures this transition in hard numbers. In the first half of 2026, Zhipu AI reported revenue of 954 million yuan, up 399.7% year-over-year. Revenue from its MaaS open platform and API business reached 825 million yuan, surging 2,735.7% YoY and accounting for 86.5% of total revenue. By contrast, last year’s corresponding share was just 15.2%, while local deployment revenue fell from 84.8% to 13.5%. Within a single year, Zhipu’s business model has fundamentally restructured.

Historically, clients purchased models for on-premise deployment, customization, and integration—typically generating one-time project revenue. Today, clients increasingly invoke cloud-hosted models directly, paying per call. This is the classic MaaS paradigm.

Moreover, Zhipu’s strategy extends beyond “lower prices for scale.” The data reveals a precise alignment: token call volume has increased over 40 times since the start of the year; the average selling price of the API has risen approximately 101%; unit token inference costs have dropped 80%; and the MaaS gross margin has improved to 24.6%. Costs are falling, revenue is expanding, usage is surging, yet average selling prices continue to climb. AI commercialization appears to be moving past early-stage price wars.

Initially, limited model capabilities forced platforms to compete on low prices or free access. But as models prove capable of executing complex tasks—code generation, tool invocation, multi-step execution—users are no longer buying a static model. They are purchasing autonomous work completion.

This evolution aligns with Zhipu AI’s management framework, which maps the business trajectory across four stages: Sell Models → Sell Calls → Sell Subscriptions → Sell End-to-End Task Results.

What Zhipu AI is ultimately selling isn’t tokens. The era of simple chatbots—where a query yields a direct answer requiring minimal tokens—is giving way to the agent era. When instructed to “build a website,” an agent must parse requirements, research, generate code, invoke tools, run tests, debug, iterate, and validate. A single task can consume hundreds of thousands to millions of tokens. Further ahead lies multi-agent orchestration, where specialized agents handle planning, search, coding, testing, and review, driving token demand even higher.

Global metrics reflect this explosion. In Q1 2026, weekly global token usage surged 250% to 22.7 trillion. OpenAI platform call volumes jumped from roughly 6 billion per minute in October 2025 to 15 billion per minute by late March 2026—a 150% increase in under six months. A Morgan Stanley report notes that top-tier LLMs are undergoing “non-linear capability leaps,” causing explosive AI adoption to collide with systemic supply bottlenecks.

The reality isn’t that tokens are becoming more expensive; it’s that demand is scaling exponentially. This dynamic mirrors the Jevons Paradox: as resource efficiency improves and unit costs decline, consumption often increases rather than decreases. AI is experiencing precisely this. Cheaper tokens empower users to delegate more tasks; smarter models encourage trust in increasingly complex workflows. Token consumption has become a self-reinforcing growth flywheel.

Zhipu AI’s Tmall storefront may appear to be selling AI quotas, but the underlying strategy is clear: transforming model capabilities into a continuously billable cloud service. In 2025, local deployment dominated revenue; by H1 2026, MaaS/API accounted for 86.5%. The company explicitly frames this progression as: “Sell Models → Sell Calls → Sell Subscriptions → Sell End-to-End Task Results.”

During the earnings call, Xiao Lei, Secretary of the Board of Directors at Zhipu AI, emphasized the significance of this structural shift: “Compared to the revenue growth rate, the shift in the revenue composition is more noteworthy.”

원문 보기