NVIDIA's Vera Rubin NVL72 computing platform is ramping into full production across 350+ factory sites in 30 countries,
Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Spanning 350+ factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand.
The Vera Rubin platform is built from chip to grid to deliver the highest performance per watt and the lowest token cost. CoreWeave's first benchmark on DeepSeek-R1 shows 10x more throughput per megawatt than Grace Blackwell NVL72, landing on the metric that matters most for power-constrained AI factories.
What makes this possible is extreme codesign across seven chips and five rack trays: Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX and Vera BlueField-4 STX, all engineered as a single system rather than assembled from separate off-the-shelf products.
The NVIDIA Vera CPU is at the center, delivering 2x single-threaded performance, 3x core-to-core bandwidth and 40% lower memory latency versus competing chiplet designs. For networking, the platform's sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency and 10x higher packet rates than off-the-shelf Ethernet. For scale-out, Spectrum-X Ethernet combines 102.4T Spectrum-6 switch systems, 1.6T ConnectX-9 SuperNICs, adaptive routing, advanced congestion control, telemetry and open operating system support, enabling 1.6x higher RDMA bandwidth than off-the-shelf Ethernet.
Leading AI infrastructure builders including CoreWeave, Microsoft, SpaceXAI and Tesla are among the first to deploy Spectrum-6 switches. NVIDIA Photonics with co-packaged optics for scale-out, the industry's first such switch in volume manufacturing, adds 5x lower power and 10x higher mean time between interruption versus pluggable transceivers.
NVIDIA's three generations of rack-scale codesign eliminated cables, fans and hoses in the compute tray, cutting assembly time from hours to one minute. A 45-degree Celsius liquid cooling inlet temperature design enables chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually.
Microsoft and Mistral announced a newly expanded partnership that brings frontier AI to Europe, combining open European models with cloud and customer-controlled environments. A new multibillion-dollar agreement focuses on expanding AI infrastructure in Europe, with Mistral adding its GPU capacity through thousands of the latest NVIDIA Vera Rubin GPUs. Compared with NVIDIA GB200 NVL72, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens.
CoreWeave ran a DeepSeek-R1 benchmark on Vera Rubin NVL72 and saw 10x improvement in tokens per second per megawatt compared with Grace Blackwell NVL72. Vera Rubin NVL72's 260 TB/s all-to-all NVLink 6 fabric removes bandwidth constraints, enabling the rack to behave as a single unified accelerator. CoreWeave deployed the NVIDIA Spectrum-X Ethernet SN6600-LD as the switching fabric for Vera Rubin NVL72, delivering 1.64 petabytes per second per rack with 100% more capacity than previous-generation air-cooled switches.
NVIDIA Vera Rubin NVL72 is powering Google Cloud's first A5X instance, now up and running for London startup Ineffable Intelligence.