Sunday, October 4, 2026
AI Infrastructure · News & Analysis
Home › Capital Markets › Report
Capital Markets · Report

AI agents now consume 5x more tokens than humans on OpenRouter, with usage growing 14x since February 2026.

Exponential growth in agent token consumption indicates rapidly accelerating compute demand trajectory beyond supervised learning.
Trade pressSlicast · October 3, 2026 at 13:10 UTC · Global · Source: Tom's Hardware
importance 72

Futurum Group CEO Daniel Newman recently highlighted a significant milestone in AI adoption: "AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher." This observation draws from an Andreessen Horowitz chart of OpenRouter data showing agents consuming 7.3 trillion tokens versus humans' 1.4 trillion tokens as of August 2026, roughly six months after agent usage first surpassed human usage on the platform.

The dominance of agent tokens, however, masks an important detail. More than 85% of those agent tokens come from cached prompts, according to a16z's analysis of OpenRouter data. This means agents are largely rereading information they have already processed, rather than handling entirely new queries.

OpenRouter, a leading AI model gateway and routing platform, categorizes each API key into one of three categories—agentic, mixed, or human—using a weighted scoring system based on factors including tool call rate, turn count, and gap timing. Since February 2026, the data reveals agents are using 14x more tokens while human usage has grown 2.8x. The mixed category, which likely represents behavior that combines both agent and human characteristics, expanded 4.7x over the same period. Depending on how that mixed traffic ultimately breaks down, agents' token advantage may vary somewhat. It's worth noting this data measures token volume rather than spending, and it comes from a single platform, though the trend does show dips in April and July.

The growth pattern extends beyond OpenRouter. McKinsey's 2026 State of AI survey found that 40% of respondents from large organizations reported scaling AI agents, up from 27% a year earlier. Similarly, a call center consultancy that tested DeepSeek on rented Nvidia H200s observed that "96% of all input was re-reading old conversation" in its own agents' September usage on Claude Code.

The reliance on cached prompts creates a new infrastructure challenge. While cached tokens cost considerably less than processing fresh prompts, they must still be held in memory. a16z notes that demand for high-bandwidth memory (HBM) is rising as models store context in the KV cache—and "the KV cache is outgrowing GPU HBM capacity." Memory constraints are already acute. Micron anticipates RAM and storage shortages will worsen in 2027 and 2028, with customers paying higher prices, while memory manufacturers prioritize HBM allocation for AI data centers. If Newman's prediction holds true and the ratio climbs from 5x to 10x, 20x, and beyond, PC buyers will increasingly compete with agents for scarce memory resources.

Read the original
AI agents now consume 5x more tokens than… · Slicast