Cerebras and OpenAI announce GPT-5.6 Sol Ultrafast Mode, delivering up to 750 tokens per second inference using Cerebras hardware at production scale.
Cerebras is now running OpenAI's flagship model at speeds no GPU cloud has publicly matched. On August 13, 2026, the wafer-scale chipmaker announced it powers GPT-5.6 Sol on a new OpenAI service tier called Ultrafast, delivering up to 750 output tokens per second and, by OpenAI's account, running the model up to 14× faster than Standard processing. Ultrafast launches first in the OpenAI API as a limited preview for a select group of customers, with access expanding as capacity grows.
The core claim is direct: frontier intelligence without the speed penalty. GPT-5.6 Sol is OpenAI's most capable model, and on Cerebras silicon it generates tokens fast enough for real-time products rather than overnight batch jobs. Both companies frame the tier as removing the tradeoff between a model intelligent enough for high-stakes work and one fast enough to use while that work unfolds.
The 750 output tokens per second figure dominates headlines, but the head-to-head comparisons Cerebras published tell more. Against output speeds reported by Artificial Analysis, the company claims GPT-5.6 Sol on Ultrafast runs 11× faster than Claude Fable 5 and 5× faster than Opus 4.8 on Fast mode.
Cerebras tested the tier against Humanity's Last Exam, a 2,500-question benchmark pitched at PhD-level difficulty. GPT-5.6 Sol on Ultrafast answered all 2,500 questions in 11 hours and 11 minutes; Claude Fable 5 required 78 hours and 27 minutes to reach comparable conclusions — nearly 7× slower by Cerebras's measurement. On GDP-Val, a benchmark for economically valuable knowledge work, the company reports a 5.6× end-to-end speedup over Standard processing with no quality degradation.
These are vendor-run evaluations. The Humanity's Last Exam comparison was benchmarked by Cerebras on July 10 and July 13–15, 2026, and the GDP-Val figure comes from its July 31, 2026 testing. Treat them as the company's own measurements, not independent results.
The mechanism explains why a relatively small chipmaker serves OpenAI's biggest model at speeds GPU incumbents have not matched. Fast inference on large models is fundamentally a data-movement problem: on GPUs, model weights shuttle repeatedly between on-chip memory and off-chip storage to generate each successive token, and memory bandwidth becomes the bottleneck.
Cerebras eliminates that movement. Its Wafer-Scale Engine packs 44 GB of SRAM onto a single wafer-sized chip, keeping model weights on-chip so tokens flow through layers pipelined across wafers without interruption. Because weights never leave the silicon, the approach scales with model size — the company's argument that the speed advantage holds as frontier models grow. For inference economics, this is decisive: the cost and latency of serving a model are dominated by how fast you can feed weights to compute, and keeping 44 GB resident on one die attacks that directly.
Ultrafast is the most visible product yet of a relationship building for months. OpenAI tapped Cerebras for $10 billion in low-latency compute earlier in 2026, and this launch puts that capacity behind the company's top model rather than a smaller or specialized one. OpenAI describes Ultrafast as "the next step" in the partnership to bring ultra-low-latency inference to its platform.
For Cerebras, the placement is significant. The startup has long argued its wafer-scale architecture is the right shape for inference even as the market's center of gravity sits with GPU suppliers — a contest playing out across the accelerator business as incumbents move to embed models directly into their silicon. Landing the serving layer for OpenAI's flagship gives Cerebras a production reference account at the market's top. The model itself anchors the GPT-5.6 family OpenAI launched, with Sol as the flagship alongside the balanced Terra and the cost-efficient Luna, so Ultrafast attaches Cerebras to the front of that lineup.
OpenAI is positioning the tier at time-sensitive, high-stakes work: incident response during active outages, financial research as market conditions shift, real-time customer support and voice, commerce, and live research loops that previously ran overnight. Early access has gone to companies across coding, commerce, and finance, including Jane Street, Podium, Basis, and Rogo.
"The increase in speed brought by Cerebras is impressive," said John Crepezzi, AI Assistants at Jane Street, in OpenAI's announcement. "It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them."
OpenAI is keeping the rollout narrow deliberately. The company says it is using the preview period to learn where an order-of-magnitude speed change creates the most value, and will expand access as capacity grows. Both companies are accepting sign-ups for updates as the preview widens.
Near-term success depends on capacity, not capability. Ultrafast is a limited preview, and both companies tie broader availability to capacity growth rather than a fixed date — so the pace of expansion is what to watch. The longer question is whether Cerebras's on-chip-memory advantage holds as OpenAI's models scale, which is exactly the bet the wafer-scale architecture is built on.