SpaceX racing toward 10 GW of AI compute capacity as part of the Stargate initiative, Elon Musk entering stated 'comfort zone' for xAI/Colossus infrastructure expansion
With the same 1 gigawatt of AI computing power, one business model might generate only tens of billions of dollars in annual revenue, while another could exceed $100 billion. What would that mean?
The answer could be a complete game-changer for the entire AI infrastructure industry.
Over the past two years, the market has measured the AI arms race by GPU counts and training cluster sizes. Whoever holds more H100s and GB200s, and whoever operates larger training clusters, is considered to have stronger AI infrastructure. But in its latest report, SemiAnalysis proposes a more radical framework: what will truly matter in the future will shift from how many GPUs you own to how many sellable tokens each megawatt of power can ultimately produce, and how much revenue that generates.
According to its Tokenomics Model and Inference Simulator, under specific assumptions for frontier models, GB300 clusters, real-world Agentic Coding workloads, and API pricing, OpenAI and Anthropic's potential annual revenue per 1 gigawatt of inference compute could exceed $100 billion. By comparison, the report estimates the annual cost of leasing an equivalent GB300 cluster at approximately $12 billion. These numbers are aggressive and highly dependent on model demand, token prices, utilization rates, latency requirements, and software and hardware efficiency—but they explain an increasingly realistic question: why have all the tech giants suddenly begun scrambling for power with such intensity?
If this prediction holds, SpaceX's ambitions may extend far beyond building rockets, launching satellites, and selling Starlink. It aims to become one of the world's largest AI compute providers. This also means that as securing compute capacity months or even a year in advance begins to generate enormous economic value, the scarcity in AI infrastructure is shifting from a competition over chips to an industrial war over power, engineering, and time. And compressing complex engineering to its limits is precisely the game Musk knows best.
The most important contribution of SemiAnalysis's report is a fresh calculation of the Revenue/MW equation. For the same GB300 cluster, if GPUs are merely rented out via the traditional model, the economic value remains primarily determined by equipment costs, depreciation, and lease pricing. But if these GPUs ultimately serve the most advanced frontier models, selling inference tokens directly through APIs, Coding Agents, Copilots, and similar products, the revenue capacity per megawatt could be an entirely different order of magnitude.
SemiAnalysis calculates that when running frontier model API inference on GB300 clusters, the potential revenue for OpenAI and Anthropic can exceed $100 million per megawatt per year—over $100 billion per gigawatt per year. This calculation incorporates variables such as model architecture, memory bandwidth, serving configuration, throughput, time-to-first-token, input tokens, cache reads, cache writes, and output tokens, while also using traces derived from real Agentic Coding workloads.
This signals a clear shift in how AI infrastructure is evaluated. In the early days, the market compared GPU counts, then FLOPS and training cluster scale. In the era of large-scale inference, the more critical metrics may gradually become tokens per second, tokens per watt, tokens per dollar, tokens per megawatt, and ultimately Revenue/MW.
This is also why the significance of GB300 goes beyond simply being "faster than GB200." According to the report's model, in the same Fable 5 inference scenario, the GB200 NVL72 corresponds to approximately $73.4 million per megawatt per year in potential revenue, while the GB300 NVL72 can increase this to approximately $99.7 million. This means that the real value created by next-generation hardware lies in producing more high-value tokens within the same power budget.
NVIDIA itself is redefining AI infrastructure with similar language. The DSX platform released in May this year has already established "token performance per megawatt" as one of the core metrics for AI Factories, extending from chips, networking, and software to power, cooling, and data center operations. As inference demand expands from ordinary chat to Agentic Coding, AI Workers, multi-agent collaboration, and continuously running tasks, this shift will become even more pronounced.
Training typically has well-defined phases and cycles, whereas inference demand more closely resembles "number of users × number of agents × token consumption per task × runtime." Once agents move from "answering questions" to "working continuously," the ceiling on token consumption expands. From this perspective, the most important variable to watch in the next phase of the AI compute war may no longer be how large training clusters can become, but rather how many tokens these compute resources can ultimately produce.