AMD is positioning distributed computing as a solution to rising costs of agentic AI inference workloads.
Tech firm AMD is urging enterprises to reconsider their computing infrastructure as agentic AI adoption drives up token consumption and cloud computing costs, pointing to distributed AI architectures as a cost management solution.
As autonomous AI agents move beyond pilot projects into departmental workflows, their continuous operation significantly increases usage-based expenses. "As AI moves from occasional prompts to continuous workloads via agentic AI, token consumption can grow quickly, where systems repeatedly reason, call tools, and iterate," said Alexey Navolokin, Asia Pacific general manager at AMD. "Each step consumes compute and system overhead, and in cloud environments, usage-based costs can accumulate."
According to Anthropic's State of AI Agents 2026 report, 57% of surveyed organizations already deploy agents for multi-stage workflows. Navolokin noted that organizations are adopting agents because they augment employees rather than simply automate jobs. "AI agents can multiply what employees are able to accomplish, rather than simply replace human work. Agents can operate continuously, handle repetitive execution and coordinate multiple steps at a scale and pace that would be difficult for an individual to sustain, while employees remain focused on decisions, creativity, domain expertise and accountability." Small businesses can use agents to support teams, creators can delegate scheduling and content distribution, and professionals can use them for research and synthesis. Within AMD, agents are already deployed in semiconductor design to accelerate issue identification and resolution.
AMD argues that enterprises may need to reconsider relying exclusively on cloud infrastructure. "Agentic AI is a distributed infrastructure challenge, with the right mix of compute needed across cloud, data center, edge and AI PCs," Navolokin said. A distributed AI architecture supports workloads across cloud, data centers, edge systems, and AI-enabled PCs because no single environment is optimal for every workload.
Organizations should practice "workload right-sizing"—determining the most efficient and cost-effective infrastructure for particular AI workloads. "Frequent or latency-sensitive workloads may benefit from local or edge compute, while larger or highly elastic workloads can continue to use the cloud," Navolokin said. Running some workloads locally on AI PCs could reduce recurring inference expenses, offsetting upfront hardware costs against avoided cloud token charges.
AMD's analysis found that at a medium workload tier of 5.7 million input tokens and 574,000 output tokens per user per day, a fleet of 500 AMD AI PCs operating under a 50% local and 50% cloud configuration could generate projected three-year savings of 40% to 60% compared with cloud-only deployments. The workload represents a knowledge worker actively using an agent harness such as Claude Code, Codex, or Hermes. AMD estimated that a fully local deployment could produce higher savings, with initial investment typically breaking even in less than 24 months.
An AMD PRO R9700 desktop configuration could support approximately 18 million tokens of AI use per day, with electricity costing about $64.80 per month. Under AMD's comparison assumptions, the three-year cost for this configuration would be $6,533 versus $81,108 for cloud-based usage. Actual costs would depend on workloads, cloud models, hardware utilization, electricity rates, and deployment configurations.
Improvements in AI PC capabilities and smaller language models are making it possible to process more AI workloads locally. "AI PCs are increasingly being viewed as AI infrastructure investments because they bring AI compute closer to where employees, data and workflows actually reside," Navolokin said. "Local systems with sufficient memory and compute can run models that support tool calling, context-rich workflows and multi-agent applications, allowing enterprises to reduce the amount of inference they need to send to cloud services. This can lower recurring cloud inference costs while also providing lower latency and greater control over where sensitive data is processed." Local processing could also strengthen AI governance as agents increasingly interact with enterprise applications and proprietary information, allowing organizations to keep sensitive workloads on-premise.
Transitioning to distributed AI architectures requires organizations to rethink their technology stacks. "The biggest challenge is balancing cost, data control, and infrastructure capability," Navolokin said. "Local AI also requires the right infrastructure, including sufficient memory, CPU and GPU capacity, security and software support." Agentic AI could also change enterprise capacity requirements, as a single employee may eventually operate multiple agents simultaneously. "One employee managing multiple agents can create much more concurrent demand for compute, memory, data access and orchestration than a traditional AI assistant," Navolokin noted. "That means enterprises should plan for balanced infrastructure across CPUs, GPUs, memory, networking and software rather than focusing on GPU performance alone."
Enterprises should build for flexibility as agentic AI technologies evolve. "An open software ecosystem can allow teams to develop locally, test at scale and deploy workloads where they make the most sense, while reducing lock-in as agentic AI evolves," Navolokin said. "At AMD, our approach is to provide an open ecosystem portfolio spanning data center, edge and AI PCs, supported by open software and standards that give customers flexibility as their AI deployments evolve."