Nvidia unveiled Rubin, an inference accelerator promising up to 10x reduction in inference costs compared to prior generations.
NVIDIA introduced Rubin, its latest AI supercomputing platform, at CES 2026. The system targets a critical market challenge: making advanced AI development and deployment significantly more accessible and cost-effective.
Rubin delivers substantial performance advantages. The platform achieves up to 10 times lower inference token costs compared to previous generations and requires only one-fourth the graphics cards that NVIDIA's Blackwell system needs to train mixture-of-experts models. The platform integrates six chips into a single system, featuring an 88-core Vera CPU and a GPU capable of up to 50 petaflops of performance.
The architecture prioritizes infrastructure efficiency through improved design and faster interconnections, directly addressing the cost barriers that have historically limited AI adoption among enterprises and cloud providers. AWS, Google Cloud, and Microsoft are positioned to benefit from these cost reductions, broadening access to advanced AI capabilities.
Rubin is scheduled to roll out to partners in the second half of 2026.