EpochAI reports quarterly AI inference costs have dropped 47%, driven by chip efficiency gains and manufacturing scale.
As artificial intelligence competition intensifies, infrastructure investments continue to grow exponentially, yet the costs of inference—the direct computational expenses incurred when AI models perform tasks—are declining at a historically unprecedented rate.
According to EpochAI, a nonprofit research institute, the cost to achieve equivalent AI performance has fallen by approximately 47% per quarter since 2023. This rate of decline means inference costs drop to one-thirteenth of their original value within a single year. While global cumulative spending on data centers is projected to exceed $30 trillion by 2050—a sum matching the scale of cumulative U.S. Treasury bond issuance—the direct computational costs of AI are accelerating downward.
EpochAI has described this cost decline as "a record-breaking pace rarely seen in human industrial history." Accounting for inflation, the rate of AI inference cost reduction is four times faster than DNA analysis technology, six times faster than computer processing power advancements, and eighteen times faster than lithium-ion battery improvements.
In concrete terms, OpenAI's January 2025 model "o3" required an average cost of 30 cents per problem to achieve a 75% accuracy rate on the "GPQA Diamond" exam, which tests doctoral-level expertise in physics, chemistry, and biology. By contrast, the "GPT-5.6 Luna" model released eighteen months later achieved identical performance on the same exam for just 0.04 cents. EpochAI compared this trajectory to a new car priced at 50 million Korean won suddenly falling to 70,000 Korean won within eighteen months.
This acceleration stems from advanced AI models and the widespread adoption of distillation training methods, which enable high-performing models to be efficiently replicated. This approach has driven costs down by 66% per quarter. As EpochAI noted: "Even if Big Tech pours massive capital and energy into developing cutting-edge models, the window to monetize technological superiority through premium pricing is fleeting."