Monday, October 12, 2026
AI Infrastructure · News & Analysis
Home › Capital Markets › Report
Capital Markets · Report

AI labs are now spending more compute on teaching models to think than on teaching them to learn, according to Startup Fortune.

Shifting compute budgets toward reasoning and post-training changes the mix of GPUs and data-center capacity that AI labs need to buy.
Trade pressSlicast · October 11, 2026 at 04:29 UTC · US · Source: Startup Fortune
importance 40

Frontier AI labs are putting more compute into post-training and inference-time reasoning, not just into ever-larger pretraining runs. That shift is why Nvidia's order books keep growing even as pretraining runs stop getting dramatically bigger.

The compute story is no longer just about making the next base-model pretraining run bigger. Epoch AI's public model tracking and later analysis of GPT-5 point to a break from the old pattern, in which each new OpenAI flagship consumed far more training compute than the last. Pretraining, the process of feeding a model the internet and hoping scale alone produces intelligence, used to dominate frontier training budgets. Now more of the money is moving elsewhere: teaching a model that already knows a lot how to reason through a problem step by step before it answers.

OpenAI's own model history shows the turning point. Epoch AI's analysis estimated that GPT-5, released in August 2025, used less training compute than GPT-4.5, OpenAI's earlier "largest and most knowledgeable" model. That does not suggest a company running short of money. It suggests a company that found another lever. Public evidence still does not reveal the full training budgets for either model, but the direction is clear: new performance gains are coming from a mix of pretraining, post-training and inference-time reasoning rather than from a simple bigger-pretraining-run story.

The shift has a rough starting date. Around September 2024, researchers at OpenAI and elsewhere began publishing techniques that made reinforcement learning on reasoning chains work at scale, rather than merely helping models sound more polished. OpenAI's o-series models were the first visible result. In January 2025, DeepSeek released R1 and made its method public: a training recipe using reinforcement learning with verifiable rewards, in which a model receives a plain correctness signal on math and code problems and learns through trial and error, without a massive human-labeled dataset. DeepSeek's paper described DeepSeek-R1-Zero developing an "aha moment," pausing mid-answer to re-check its own logic, a behavior no one had explicitly trained into it.

That mattered because the approach was cheap and copyable. According to multiple AI researchers tracking the field through 2025, including the widely read "State of LLMs" roundup by researcher Sebastian Raschka, nearly every major lab, open-weight and closed, shipped a reasoning variant of its flagship model within months of R1's release. Anthropic's Claude models gained extended reasoning modes, and Google built reasoning into Gemini. Once DeepSeek showed the recipe worked outside a frontier lab's walls, few had reason to skip it.

This matters most for anyone still modeling AI infrastructure demand from pretraining FLOPs alone. A reasoning model doesn't just train differently; it runs differently. Each time it answers a question, it generates a chain of hidden reasoning tokens before producing the visible response. Industry estimates cited across infrastructure research put this at 10 to 100 times more tokens per query than a non-reasoning model performing the same task. That compute cost is not paid once during training. It is incurred every time a user hits send.

This is the mechanism linking the shift in compute allocation to Micron's financial results. Micron announced record fiscal fourth-quarter revenue of $54.23 billion on September 30, with full-year revenue reaching $133.19 billion. The company and analysts have tied these figures directly to surging demand for high-bandwidth memory (HBM), which feeds AI inference clusters. TrendForce, the Taiwan-based research firm, estimates that 2026 capital expenditure by the nine largest cloud providers will exceed $886.7 billion, with spending extending beyond GPUs into AI servers, liquid cooling, advanced packaging, high-speed interconnects, power infrastructure and memory. Separately, Counterpoint Research projects that HBM demand for custom AI processors will grow 35 times between 2024 and 2028.

This reframes the AI capex debate that bond investors and chip analysts have been having. The concern has been that pretraining scaling is hitting diminishing returns, so eventually GPU orders would stop. That concern assumed GPUs were mainly for training. They are increasingly used to run models after training, and a reasoning model that consumes around 50 times the tokens per answer needs that inference capacity whether the next pretraining run is twice as large or barely larger. Nvidia CEO Jensen Huang has made exactly this case publicly, arguing that reasoning models will keep pushing GPU demand higher even as pretraining scaling slows, a claim covered by Computer Weekly, among others.

There is a second-order effect as well, and it concerns who wins the next round of benchmarks. If the bottleneck has shifted from "who can afford the biggest pretraining cluster" to "who has the better reinforcement learning recipe," the competitive math changes. A lab with a strong RL approach and a mid-sized model can potentially out-reason a much larger model that only benefited from pretraining scale. DeepSeek strengthened this argument by releasing a competitive open-weight reasoning model with a training recipe the rest of the market could study. It is why OpenAI, Anthropic and Google are now competing on reasoning benchmarks rather than parameter counts, and why the next leap in AI capability is more likely to appear in a training recipe paper than in a bigger data center groundbreaking.

None of this means pretraining disappears. Someone still has to build the base models that reasoning techniques are layered onto. But the compute math has flipped, and the chip demand story built on top of it now depends far less on pretraining ever scaling up again.

Read the original
AI labs are now spending more compute on… · Slicast