AI 추론 유닛 가격이 처음으로 1달러 미만으로 하락하며, 골드만삭스는 신규 컴퓨팅 공급 과잉이 하이퍼스케일러의 자본지출 기반을 위협한다고 경고했다.
AI inference service unit pricing fell below $1 per million tokens for the first time in August. This breach of a critical psychological threshold fundamentally undermines the core narrative that has driven AI sector valuations over the past two years. In a recent report, Rich Privorotsky, head of Goldman Sachs’ One-Delta trading desk, noted that the Silicon Data LLM Token Expenditure Index (SDLLMTK)—which tracks the usage-weighted average price per million tokens—dropped 29% in August alone. It closed the month at approximately $0.97, marking an all-time low and a cumulative decline of more than 50% from its May peak of roughly $2.05.
This price collapse does not signal shrinking demand; rather, it reflects a deteriorating unit economics model. According to a JPMorgan data center report, OpenRouter’s token usage grew approximately 47% month-over-month in August, while dollar spending increased by only about 7%. The widening gap between rising volume and falling prices reveals a harsh reality: equity markets price returns in dollars, and hyperscaler capital expenditures are calibrated against dollar-denominated revenue expectations.
Privorotsky stated bluntly that combining higher demand with lower prices does not automatically constitute a positive outcome when the pace of price declines outstrips consumption growth. He emphasized that August’s data shattered, for the first time, the foundational assumption that “more tokens equals more revenue”—a premise that has anchored AI equity pricing for the past two years.
**Three Premises Collapsing Simultaneously**
The sustained decline in token prices is not accidental. Privorotsky expressed skepticism toward per-token billing as a durable business model, noting that cloud inference pricing depends on three premises holding simultaneously: models being too large to run locally, users lacking viable open-source alternatives, and workloads being sufficiently bursty to make in-house compute uneconomical. All three are now unraveling in tandem.
On the hardware front, devices such as RTX Spark-class laptops, DGX Spark workstations, Mac Studios capable of running 70-billion-parameter models, and NPUs delivering 40 to 75+ TOPS have compressed the marginal cost of processing large token volumes to near electricity-cost levels. Once an enterprise’s monthly API bill exceeds the amortized cost of a $5,000 to $15,000 device, it transitions from a token consumer to a one-time hardware buyer, abandoning the recurring software subscription model that historically yielded 40% margins.
On the model competition front, Meta began rolling out Muse Spark 1.3 on September 2. The model already ranks alongside GPT-5.6 Sol and Claude Opus 5 on independent benchmarks, demonstrating standout performance in agentic tasks and code generation. Privorotsky noted that frontier advantage is now measured in weeks rather than years—it is not a sustainable moat, merely a product cycle. Whenever the next-best model is sufficiently close and cheaper, enterprises route around premium tiers, and the token price index serves as the aggregate expression of those routing decisions.
Silicon Data itself has cautioned that the index decline may stem from list-price reductions, user migration to lower-cost open-source models, or both, and does not imply that overall AI usage is contracting. Yet the underlying issue remains: OpenRouter routing volume has surged while H100 GPU rental prices are falling, directly contradicting the prevailing “token demand explosion” narrative.
**Credit Markets Flash Warning First**
The Goldman Sachs report underscores that this AI capital expenditure cycle consists of locked-in, asset-heavy commitments rather than flexible operating expenses. The five rated hyperscalers are projected to spend approximately $737 billion in combined 2026 capital expenditures, representing roughly 38% of their revenue. Moody’s has already issued warnings regarding free cash flow compression, the balance-sheet shift from asset-light to asset-heavy structures, and lease commitments that, while not appearing as traditional bonds, substantially constrain issuers.
Credit markets have reacted ahead of equities. According to a tally of relevant bonds, 78 of the 91 hyperscaler bonds issued in 2026 had fallen below their issuance price by the end of August. Privorotsky characterized this as a “credit repricing,” emphasizing that equity valuation multiples remain a lagging indicator.
The transmission chain is clear and consequential: if token prices fall 30% while usage grows 20%, inference revenue declines. As inference revenue contracts, depreciation and interest costs from the 2025–2027 infrastructure buildout continue to rise, causing return on invested capital (ROIC) to collapse. Once ROIC deteriorates, the market does not require a dramatic “bubble burst” narrative; it simply adjusts valuation multiples to reflect a utility-like enterprise burdened with substantial assets but constrained pricing power.
**GPT-6 Astra: The Only Remaining Narrative-Reversal Variable**
Amid the current outlook, Privorotsky identifies OpenAI’s next-generation model, Astra, as the only near-term catalyst capable of reversing the trend. On September 1, OpenAI announced that Astra has reached a critical cybersecurity threshold under its Preparedness Framework, becoming the first model classified in that category.
However, Goldman Sachs’ trading desk remains cautious. A highly restricted model subject to access controls, monitoring requirements, and higher usage friction does not automatically replenish the token revenue pool. If high-value workloads remain confined to controlled testing environments without entering the public billing system, Astra could actually reduce the total number of billable tokens.
The report concludes that Astra represents the final near-term variable capable of tilting the revenue mix back toward high-value tiers. Until then, August’s token pricing data stands as the most honest signal of current market conditions. Demand can grow indefinitely, but if the price at which that demand arrives cannot cover the cost of debt issued to finance data centers, equity valuations will still face downward pressure.
Notably, Wall Street’s sentiment toward Meta remains strongly positive. Over the past three months, analysts have assigned the stock 37 buy ratings, 6 holds, and 0 sells, resulting in a consensus “Strong Buy.” The average price target of $753.08 implies approximately 22.2% upside. Nevertheless, the core message from Goldman Sachs’ trading desk is unequivocal: stock-level optimism cannot obscure industry-wide structural pressures. The pace of compute supply expansion may be outstripping demand growth, and the financial cost of that imbalance will ultimately materialize in degraded capital expenditure returns.