Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomePower & EnergyReport
Power & Energy · Report

Analysis emphasizes that electrical grid capacity remains the primary bottleneck constraining AI model scaling regardless of algorithmic efficiency.

Reinforces that power infrastructure investment and grid interconnection timelines dictate actual AI compute deployment speeds.
Trade pressSlicast · August 26, 2026 · US · Source: Google News
importance 65

For three years, the contest for artificial intelligence (AI) has been scored one way: whoever trains the largest frontier models on the most advanced chips in the largest data centers wins. That premise drives US export controls, which aim to deny China the hardware required to build ever-larger systems. It also drives the infrastructure buildout now straining power grids from Virginia to Texas, where firm electricity, not capital, has become the binding constraint on how much compute comes online. Underpinning this dynamic is an assumption worth examining: that the frontier model in the cloud is where the game is decided.

A study published in November 2025 by researchers at Stanford University and Together AI, an American AI infrastructure firm, casts doubt on that assumption. The authors propose a single yardstick for AI efficiency: intelligence per watt, defined as the task accuracy a system delivers for each unit of power it consumes. Running more than a million real-world queries across 20-odd compact models and eight types of hardware, they found that models running locally on consumer-grade silicon—the kind found in modern laptops—correctly answered 88.7 percent of single-turn chat and reasoning queries. Between 2023 and 2025, intelligence per watt improved 5.3 times, and the share of queries a local model could handle climbed from 23 percent to 71 percent.

This trend is mirrored by another. The price of intelligence has been falling at a comparable pace. OpenAI’s flagship model cost $30 per million input tokens when GPT-4 launched in early 2023; a little over a year later, GPT-4o performed the same work for $2.50, and a cheaper variant matching the original on most tasks arrived at a fraction of that cost. The point is not the exact figure. It is the slope. Efficiency is the fastest-moving component of this ecosystem, advancing more rapidly than any infrastructure built from steel and copper.

None of this renders the cloud obsolete. The most complex queries will continue to route to frontier models for years to come. Yet this shift reframes the actual stakes of the competition, pulling focus directly to the energy dynamics that now dictate AI’s trajectory.

Power demand from AI has typically been treated as a hurdle to overcome: expand grids, queue for gas turbines, license reactors—all so larger models can run in larger facilities. That is the logic driving the current data-center boom, and it explains why interconnection queues in the largest US markets now stretch five to seven years. Compute is downstream of watts, and watts are downstream of infrastructure that takes a decade to permit and build. In a working analysis I have been developing on the global data-center buildout, the recurring finding is that the United States and the European Union (EU) will underdeliver against their own 2030 power targets, while China, coordinating generation and grid through the state, will not. Energy, not chips, is where that race will be won or lost.

Intelligence per watt flips this logic. If a growing share of demand can be served by efficient hardware deployed closer to the user, individual queries no longer require additional gigawatts of centralized, firm capacity. Efficiency—how much useful inference a system extracts from each watt—begins to matter as much as raw scale. Unlike transmission lines, efficiency is a rapid variable. It increased more than fivefold in two years, outpacing the expansion of physical grids.

For those who have watched energy emerge as the true ceiling on AI ambition, this represents a more consequential frontier. Durable advantage may belong less to those who assemble the largest frontier compute clusters and more to those who convert energy into useful output with maximum efficiency—a contest determined by chip design, manufacturing, and system architecture rather than by the sheer volume of capital invested in data centers.

US export controls operate on a specific theory of leverage: deny China access to the most advanced accelerators, and you constrain its ability to build frontier-scale systems. This logic holds only as long as capability remains tightly coupled to frontier scale.

Pushing most everyday demand onto efficient hardware erodes that leverage at the margin. A critical distinction must be held here: the Stanford findings apply to small models running on consumer-grade silicon—laptops, not server farms. China’s parallel strategy focuses on scaling domestic inference chips. Huawei’s Ascend series, which developers’ own tests place at roughly 60 percent of an Nvidia H100’s inference performance, has already become the default for a growing share of Chinese deployments despite trailing the frontier. A competitor barred from the highest-end chips, yet competitive in efficient, readily deployable hardware across multiple tiers, has a viable path to meeting the bulk of real-world AI demand without ever reaching frontier parity.

This is not an argument that export controls are futile. The chokepoint remains highly effective precisely where intended: frontier training, the largest models, and the most advanced semiconductor nodes. Rather, it highlights that export controls address only one layer of the stack, leaving the efficiency layer—where denial is less effective and a manufacturing-intensive competitor holds comparative strength—largely unaddressed.

Intelligence per watt remains on the fringes of strategic discourse for a simple reason: it is difficult to quantify. Analysts and policymakers currently benchmark AI dominance using easily tallied metrics—chips shipped, parameters trained, data-center megawatts announced. Measuring watts of useful inference resists such straightforward accounting.

This observation echoes a longstanding principle of energy geopolitics: the metrics used to measure a competition dictate the strategies employed to win it. Track barrels, and nations pursue reserves. Track firm capacity, and they prioritize grid development. A frontier-and-chip scoreboard fuels a race to construct the largest systems and restrict rivals’ access to matching hardware. A watts-based scoreboard would instead reward grid reliability, efficient silicon, and the industrial capacity to deploy that silicon broadly—a fundamentally different competition, one in which the current US posture is not clearly dominant.

Adopting such a scoreboard would necessitate corresponding policy shifts. It would require coupling chip export controls with neglected priorities: treating grid capacity, transmission infrastructure, and firm power supply as core components of AI strategy rather than deferred maintenance issues, and measuring progress by useful inference delivered per watt rather than accelerators shipped. Export controls buy time at the frontier, but they do not construct the power infrastructure that ultimately determines who can deploy AI at scale. That infrastructure remains largely left to market forces in Washington.

The Stanford and Together AI paper is a single study, and its authors explicitly acknowledge its limitations: reliance on single-turn queries, idealized routing assumptions, and the continued necessity of frontier models for the most complex tasks. It does not definitively resolve how AI capacity should be measured. Nevertheless, it places the issue squarely on the agenda. For a policy framework built almost entirely on opposing assumptions, this warrants serious consideration before the next iteration of export controls is drafted.

Read the original
Analysis emphasizes that electrical grid… · Slicast