Friday, October 2, 2026
AI 인프라 · 뉴스 & 분석
홈 › 반도체·하드웨어 › 리포트
반도체·하드웨어 · 리포트

Cerebras demonstrates a 5x throughput gain in AI inference through its disaggregated chip architecture.

Improved inference efficiency via disaggregated design could reduce per-inference compute costs and accelerate adoption of alternative chip architectures.
업계 전문지Slicast · 2026년 10월 1일 20:19 UTC · 미국 · 출처: Unite.AI
중요도 70

Cerebras Systems announced October 1, 2026 that it achieved 5x inference throughput improvement through disaggregation—splitting inference into separate hardware pools—without losing token generation speeds or requiring additional Cerebras systems. The company published its findings in "Disaggregated Inference From the Ground Up," a blog post by Isaac Tai and Zhenwei Gao that opens a planned series exploring the technique.

The post frames disaggregation for readers who have encountered the concept but seek practical grounding, particularly the claim that prefill is compute-bound while decode is memory-bound. It builds an explanation from first principles, starting with how accelerators balance arithmetic against data movement.

**Prefill and Decode Impose Different Hardware Demands**

Arithmetic intensity—the ratio of floating-point operations to bytes transferred between memory and compute units—varies significantly between inference phases. Adding two matrices performs 1 FLOP for every 6 bytes moved, yielding 0.167 FLOP per byte; this ratio remains constant as matrices scale. Matrix multiplication differs fundamentally: each output value draws from an entire row of one input and an entire column of the other, so loaded values contribute to multiple outputs and arithmetic intensity grows with input size.

Inference consists of chained matrix multiplications between fixed model weights and input tokens, executing in two distinct phases with different intensity profiles. Prefill processes the entire prompt in parallel as a large matrix operation—real-world prompts can span thousands or hundreds of thousands of tokens. Decode generates tokens one at a time, relying on the KV cache, which stores keys and values computed for earlier tokens to avoid recomputation.

Both phases must move the complete model weights—potentially hundreds of gigabytes or terabytes—to compute units for every token generated. Memory movement remains nearly constant while arithmetic intensity drops after prefill, a gap that explains why adding raw compute capacity does not necessarily accelerate token generation during decode.

**Disaggregation Separates Inference Into Independent Pools**

In production, inference servers handle many requests concurrently, and when prefill and decode run on shared hardware, compute-intensive prefill can stall active decode requests. Schedulers must navigate competing priorities: delivering new requests to their first token quickly, maintaining smooth response streaming, and maximizing total throughput. Batching lets concurrent requests share the work of reading model weights, but larger batches can lengthen each decode step, raising total throughput while slowing individual user token delivery.

Disaggregation is a systems design pattern that runs prefill and decode in separate hardware pools. Once separated, operators can allocate hardware, set batching policies, and optimize each stage independently for latency or throughput: systems targeting strict time-to-first-token requirements can reserve more prefill capacity, while those designed for smooth streaming can dedicate a larger or more tightly scheduled decode pool. Each pool can scale individually.

Separation introduces operational complexity. After prefill builds the KV cache, that request-specific state must transfer to the decode pool and load into memory before generation resumes; model weights are already loaded in both pools. The handoff incurs network and coordination overhead, either pool can sit idle if capacities mismatch demand, and added latency depends on whether the cache moves across colocated machines or across regions. Disaggregation proves most compelling at scale, where gains from independently sizing and scheduling pools outweigh transfer and operating costs, and it fundamentally changes the serving system's control interface rather than merely smoothing streaming.

**Heterogeneous Hardware and Early Results**

Cerebras is leading development of heterogeneous disaggregation, combining multiple chip types in one inference system and assigning memory-bound or compute-bound tasks to appropriate hardware. The company contrasts its wafer-scale design, which distributes SRAM alongside compute across the entire wafer, with GPUs, which stage model data from high-bandwidth memory through smaller on-chip caches.

A peak memory-bandwidth chart dated September 10, 2026 lists Cerebras WSE-3 on-chip SRAM at 21,000 TB/s per wafer, an unnamed on-chip SRAM accelerator at 150 TB/s, and HBM4 GPUs at 23.3 and 22 TB/s. The chart notes that SRAM figures sum local memory bandwidth across a processor while HBM figures measure off-chip traffic, describing different memory tiers rather than direct token-speed comparisons.

A separate benchmark from September 10, 2026 tested GPT-oss-120B with high reasoning on 10,000 input tokens. Results showed Cerebras at 1,669 output tokens per second, SambaNova at 708, Groq at 475, Microsoft Azure at 319, Nebius at 294, and Baseten at 293.

Cerebras reported that in traditional aggregated systems, capacity gains require deploying additional hardware. By partnering with accelerators to handle prompt processing, it increased capacity 5x in early tests using the same WSE footprint. The company has announced partnerships with multiple hardware partners to accelerate token delivery; a diagram shows AWS Trainium and AMD Helios Instinct GPU systems among the prefill options feeding a Cerebras decode pool.

Agentic applications represent a compelling disaggregation use case, given their reliance on long, multi-turn workflows where context grows across model calls and latency at each step compounds. Cerebras will cover the hardware and software stacks involved and the economic trade-offs of deploying disaggregated inference at scale in forthcoming installments.

원문 보기
Cerebras demonstrates a 5x throughput gain in… · Slicast