Analysis of agentic AI token economics highlights how inference-heavy workloads are reshaping compute demand, with specialized inference-optimized chip companies Groq, Cerebras, and SambaNova emerging as key beneficiaries.
"AI tokens are working capital," said Tarun Pant, Director of Customer Engineering at Google Cloud, while moderating a discussion on startup AI economics at the Google Cloud Startup Summit 2026 in Bengaluru.
The observation captured a problem recurring throughout the event. Discussion centered on smaller models, caching, model routing, observability and cost per business outcome—factors becoming central as AI products move beyond demos and begin serving users at scale.
The core tension is this: an agent may retrieve company data, reason through several steps, invoke software tools, call another model and retry an action before producing a result. Voice and video add further processing overhead. A token can become cheaper while the complete workflow becomes more expensive.
Gartner expects AI inference costs per agentic workflow to increase more than fivefold through 2028. Its argument: falling model prices are being offset by more complex workflows that consume more reasoning and more tokens. The FinOps Foundation is seeing the cost problem spread quickly. Its 2026 State of FinOps survey covered 1,192 respondents representing more than $83 billion in annual cloud spend. Ninety-eight per cent said they now manage AI spend, compared with 63 per cent in 2025 and 31 per cent in 2024. Visibility into AI costs, allocating those costs to business units and measuring returns remain recurring problems.
**The New AI Cost Stack**
The price of the model is only one part of what a production AI workflow costs. Context—how much data must be sent to the model for each request—is the first variable. The model itself is the second: does the task need a frontier LLM or a smaller specialised model? Tools come next: how many retrievals, APIs and reasoning steps does the workflow invoke? Reuse—whether prompts, context or responses can be cached instead of recomputed—is the fourth factor. Finally, outcome: what business result did that total compute spend produce?
**The P&L Has Entered The AI Conversation**
Amit Kumar, Managing Director of Digital Natives Business at Google Cloud India, set the tone in his opening address. "The era of adopting AI for novelty is over," he said.
Google's pitch to startups was direct: stop treating AI capability itself as success and connect deployments to financial results. Kumar divided that value into cost optimisation, customer retention and revenue generation. The broader recommendation was to embed autonomous AI into business processes while establishing a measurable route to ROI before expanding deployments.
The timing is significant. AI startups spent the first phase of the generative AI boom securing access to capable models and sufficient compute. They are now dealing with what happens when successful products generate a recurring inference bill. Google itself acknowledged this problem in August when it introduced additional cost controls for agent workloads, including expanded billing options and tools intended to give companies better visibility over AI expenditure. A July update added early anomaly detection for AI services and spend caps to Google Cloud Budgets.
**Three Tests Before An AI Agent Scales**
Moving an agent from a controlled demo into production changes the test from capability to usability, control and economics. Operational adoption asks: will the people doing the work actually use it without learning a new technical language? Enterprise governance asks: can data access, permissions and actions remain controlled and auditable? Unit economics asks: does each successful outcome create enough value to justify its inference and infrastructure cost?
**Cost Per Token Can Hide The Real Bill**
The panel discussion "Unpacking Startup AI Economics & Real ROI" decoded the issue down to engineering choices. Rachit Parekh, Partner at Accel, spoke about optimising token economics. Mohit Sharma, Co-Founder and Head of Technology at Kutumb, discussed an application serving a few million users who send substantial volumes of user-generated audio and video. For a service operating at that volume, sending every task to the largest available language model quickly becomes an expensive default.
The discussion covered fine-tuned smaller language models, LLMs, caching and choosing different models for different jobs. The metric that kept returning was cost per business outcome rather than cost per token.
Syed Shahrukh Ahmad, Co-Founder and Product Owner at BeVigil, CloudSEK, approached the same problem through security data. Large volumes of threat information must pass through an AI compute pipeline, and different tasks do not necessarily need the same model. LLM observability becomes part of cost control because a company cannot reduce what it cannot see: which model was called, how much context went in, how many calls followed, whether the response was cached and whether the task succeeded.
That shift is already visible outside the conference room. Google told startups in a September 28 post that teams scaling more sustainably were moving away from a one-size-fits-all model strategy, combining frontier APIs with smaller models according to the workload. Stanford's 2026 AI Index noted another reason to expect more model mixing. As of March 2026, Anthropic, xAI, Google and OpenAI were separated by only 22 Elo rating points on the Arena leaderboard, a system based on head-to-head user preferences. Stanford said that convergence is moving competitive pressure towards cost, reliability and domain performance.
**Building Around More Than A Model**
That same approach appeared in a conversation between Amit Kumar and Samir Shah, Vice President of Applied Science for Consumer Shopping at Flipkart. Shah described the rapid churn in foundation models and argued against putting too much engineering investment into any single one. His emphasis was on the parts around the model: tools, entities, prompting, knowledge and the customer experience. Efficient entity and tool design also affect token use, he said. Poorly designed tools can push more unnecessary context through a model and raise the cost of each interaction.
For Flipkart's conversational shopping work, this is especially relevant. A person shopping for clothes for a wedding may not type the neat product keywords that conventional search was built around. The user may describe the occasion, mix Hindi and English, refine the request during the conversation and only later settle on a product. Flipkart's AI Mode can be entered directly, Shah said, while some search queries can also be shifted into a conversational experience when the system determines that dialogue would better serve the user's intent.
His voice example was specific to India. Customers may say "uh-huh" or "okay" while another person is speaking without intending to take over the conversation. Voice activity detection systems trained around different conversational habits can interpret those signals as interruptions. For a shopping bot, that mistake is immediately visible: it may cut itself off whenever the customer simply says "uh-huh" or "okay".
**Context Also Has A Compute Cost**
Parul Sahoo, Customer Engineer at Google Cloud, put three conditions around a fictional warehouse automation deployment: frontline adoption, enterprise governance and unit economics. Her demo used three personas at the fictitious Simple Logistics to show what happens when an agent leaves a sandbox and is allowed to interact with operations.
The cost example was revealing. An early implementation was described as pushing thousands of uncompressed PDF documents in