AMD and Cerebras announce technical partnership for ultra-low-latency AI inference, combining Helios rack-scale infrastructure with Cerebras' Wafer-Scale Engine (WSE).
AMD and Cerebras Systems have announced a technical partnership to deliver a disaggregated AI inference solution combining AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine. Unveiled at Advancing AI 2026, the solution is designed to deliver ultra-low latency for advanced AI applications while dramatically increasing throughput and efficiency. The two compute engines are expected to deliver up to 5x higher tokens per second per watt.
AI inference workloads increasingly have different requirements across latency, throughput, token capacity, cost, and scale. High-volume workloads prioritize maximizing token generation, while coding, real-time copilots, live agents, and agentic workflows demand faster response times. This diversity is driving demand for heterogeneous infrastructure that matches compute technologies to specific workload requirements.
The AMD and Cerebras solution addresses this challenge through disaggregated inference, optimizing the two primary stages of the workflow independently. AMD Helios provides ultra-high throughput, processing prompts and large context windows. The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation with ultra-low latency. By connecting these engines through one integrated workflow, the companies are creating a differentiated platform for ultra-low-latency inference without sacrificing throughput or scale.
AMD Helios provides the high-throughput prompt engine, rack-scale efficiency, and deployment scale required to process large numbers of complex requests. Cerebras Wafer-Scale Engine technology provides the ultra-low-latency and decode performance needed to return tokens in real time. The solution is designed specifically for the ultra-low-latency segment of the inference market, with AMD Helios as the foundation for high-throughput and balanced inference workloads across the data center.
"AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," said Dr. Lisa Su, chair and CEO of AMD. "AMD Helios delivers leadership performance and scale for the broadest range of inference workloads. Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI."
"The demand for ultra-fast inference is growing at an unprecedented pace. Cerebras delivers the world's fastest, ultra-low-latency inference," said Andrew Feldman, CEO and co-founder of Cerebras. "Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers."
Fast token generation is becoming increasingly important as AI moves into software development, autonomous agents, robotics, scientific discovery, and other applications where response time directly shapes user experience and system usefulness. Cerebras plans to deploy AMD Helios systems in its data centers, with the joint solution expected to become available initially through Cerebras Cloud in the second half of 2026.