NVIDIA Blackwell GPUs are powering OpenAI's GPT-6 Astra Ultrafast model, now available in the OpenAI API and to eligible ChatGPT Work and Codex users in production.
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is now available through the OpenAI API and to eligible ChatGPT Work and Codex users.
Leveraging inference optimizations built into OpenAI's models to tap NVIDIA Blackwell architecture, Ultrafast delivers up to 8x faster token generation than Astra Standard mode. For developers, accelerated generation shortens coding agents' edit-test-debug cycles, reduces latency between tool calls, and makes interactive applications feel more responsive.
The performance gains matter most when applied repeatedly across workflows—an agent writes code, invokes a tool, checks the result, and decides on the next step. Ultrafast brings Astra's full capabilities into these time-sensitive loops, enabling NVIDIA infrastructure to deliver useful model outputs when developers need them.
"NVIDIA's deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs," said Philippe Tillet, inference lead at OpenAI. "Astra can turn that knowledge into high-performance kernels that make NVIDIA hardware compelling across the full frontier of latency, throughput and cost. With Astra Ultrafast, that means faster model responses as agents write code, use tools and work through complex tasks."
OpenAI continues to refine inference performance after deployment by using its own models to optimize inference software on NVIDIA GPUs, taking advantage of the platform's programmability to test and implement improvements. This iterative approach accelerates model responses and increases deployed infrastructure productivity over time.
"Our work with NVIDIA is helping us make AI faster and more useful," said Uday Ruddarraju, chief technology officer of compute at OpenAI. "We used our internal models to optimize inference on NVIDIA GPUs, and NVIDIA's programmability helped us deliver the acceleration behind Astra Ultrafast."
A programmable NVIDIA platform enables developers and researchers to reuse infrastructure across training, inference, and reinforcement learning as models evolve, improving resource utilization and avoiding overprovisioning for individual workloads. Developers can access GPT-6 Astra Ultrafast through the API today; see the Ultrafast guide for access, pricing, and implementation details.