Two neocloud operators selected Cerebras' multi-silicon inference solution within a one-week period, marking rapid adoption of its wafer-scale compute architecture.
Gimlet Labs and Cerebras Systems announced on September 28 a collaboration to deliver ultrafast AI inference through Gimlet Cloud, a purpose-built disaggregated inference cloud spanning datacenter infrastructure to developer APIs. The companies plan speeds of up to 3,000 tokens per second for demanding agentic and real-time applications, with the first Cerebras-powered Gimlet Cloud datacenter expected to come online in late 2026. Gimlet is a launch partner for the Cerebras CS-4, plans to deploy 100 megawatts of Cerebras-powered inference capacity, and expects CS-4 access in 2027.
"Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets," said Zain Asgar, Co-Founder and CEO of Gimlet Labs.
The announcement set the stage for a parallel move. General Compute announced a multi-year agreement with Cerebras to deploy ultra-fast inference for agentic coding workloads, with availability planned for Q1 2027 and the buildout financed by a $400 million debt facility from Upper90. "Agentic coding is the clearest example," said Finn Puklowski, Co-Founder and CEO of General Compute. "Agents make thousands of sequential calls, and latency compounds into wall-clock time." Gimlet Labs, backed by Andreessen Horowitz and Menlo Ventures, plans to expand the collaboration across software, APIs, developer tooling, and production operations.
Two clouds selecting Cerebras within one week represents the first commercial test of whether multi-silicon inference can beat homogeneous GPU clouds on speed at production scale. Gimlet's technical approach is the more detailed commitment. Gimlet Cloud combines the Cerebras Wafer Scale Engine with GPUs in an integrated solution, using inference disaggregation to run each phase of model execution on the silicon best suited to it. Cerebras has already posted 4,400 tokens per second per user on GPT-OSS-120B according to Artificial Analysis, but that per-user benchmark measured at low concurrency says little about serving frontier models to thousands of simultaneous agents.
Gimlet's answer is disaggregation: keep decode on SRAM, prefill on GPUs, and build the data center around that split from the start. Enterprise demand justifies this architecture. In an ETR AI Product Series survey from September 2026, 69.9% of respondents named performance among the most important factors when assessing a foundation model (n=511), ahead of cost at 68.7% and up from 63.8% in March 2025. Among a separate cohort, 78.8% named efficiency improvements the primary KPI for judging their own AI applications (n=600). These figures suggest that inference speed is graduating from a benchmark category into a distinct market with its own purpose-built clouds.
The CS-4 provides the missing infrastructure. At Hot Chips 2026, every major chip team converged on memory efficiency and synchronization-penalty elimination as the central design battleground. The CS-4's rack-level innovations—doubled per-wafer power delivery, direct liquid cooling, and a networking overhaul—make Gimlet's model buildable. In architectural detail, the CS-4 doubles off-wafer bandwidth to 2.4 Tb/s per wafer and 7.2 Tb/s per system, speaks standard RoCE v2 RDMA over Ethernet, and adds Direct Wafer Links joining wafers at 2 microseconds. Arista Networks Etherlink switches scale the fabric across racks. This standard-Ethernet approach allows a neocloud fielding NVIDIA GPUs, AMD hardware, and other accelerators to integrate a wafer-scale decode tier into a single orchestration layer without proprietary interconnect islands.
The 3,000 tokens-per-second target requires stacking disaggregation techniques across hardware. Prefill-decode disaggregation runs prompt processing on GPUs and hands token generation to the Cerebras wafer, whose 44 GB of on-wafer SRAM per wafer moves data at 43.2 PB/s across three WSE-3 Turbo processors in a CS-4 rack, eliminating the HBM bottleneck capping GPU decode speed. Attention-FFN disaggregation splits within each model layer, placing attention on GPUs and expert activation on SRAM-centric chips, trading some latency for higher throughput. Speculative decoding runs a draft model at extreme speed on SRAM hardware while GPUs verify large batches of tokens efficiently. The CS-4 already achieves more than 1,000 tokens per second on models exceeding 10 trillion parameters, and speculative decoding multiplies effective decode speed by two to three times when the draft model runs fast enough. Gimlet claims the combined approach yields 3 to 10x higher interactivity at a given throughput target, or 3 to 10x more throughput per kilowatt at given interactivity. The 3,000 tokens per second number remains a forward-looking plan rather than a measured result, plausible in architecture but requiring production validation.
Fine-grained disaggregation requires heterogeneous hardware in the same building on a fabric everything can address. Attention-FFN disaggregation ping-pongs between GPU and wafer at every layer, making cross-datacenter links disqualifying, and even prefill-decode splits degrade over distance. The CS-4's Ethernet fabric makes this possible. Cerebras' decision to trade proprietary isolation for Ethernet interoperability has produced its first purpose-built cloud customers.
The scale commitments warrant equal scrutiny with speed claims. Gimlet plans 100 megawatts of Cerebras-powered inference capacity, implying on the order of 700 to 800 racks—a material allocation from a supplier whose disclosed priorities include OpenAI, G42, and AWS. Cerebras contracted a manufacturing expansion exceeding 10x in 2026 across three contract manufacturers and reported $25.4 billion in remaining performance obligations at Q2, suggesting capacity exists on paper. Whether a venture-backed neocloud secures 2027 CS-4 allocation alongside Master Relationship Agreement customers remains a supply chain question rather than a software one. The sequencing also reveals a gap: the first Cerebras-powered Gimlet Cloud datacenter is expected in late 2026, while CS-4 access arrives in 2027, indicating the initial deployment will run current-generation systems at lower ceilings than the headline target. For Cerebras, the deal diversifies a customer base concentration-risk disclosures show is narrow and converts the CS-4's Ethernet openness into a distribution channel.