CoreWeave leads deployment of Nvidia's next-generation Vera Rubin NVL72 platform
CoreWeave announced in a press release on June 1 that it completed the industry's first build and system-level validation of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud. According to CoreWeave's announcement, the NVL72 rack contains 72 GPUs and 36 CPUs, equipped with 260 TB/s sixth-generation interconnect fabric; the company stated that the platform targets large-scale inference, agentic AI, and continuous inference workloads. The press release and CoreWeave blog attributed the performance and efficiency improvements to rack-scale design and cited comments from Jane Street's Craig Falls regarding improved iteration speed. DatacenterDynamics and SiliconANGLE complemented the coverage, citing Michael Dell's LinkedIn confirmation and theCUBE Research analysts' comments on collaborative engineering between cloud providers, platform operators, and infrastructure suppliers as agentic workloads expand.
CoreWeave announced in a press release on June 1 that it completed the industry's first build and system-level validation of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud. The company's documentation and blog articles indicate that the NVL72 rack integrates 72 GPUs and 36 CPUs and leverages 260 TB/s sixth-generation interconnect fabric to achieve rack-scale connectivity bandwidth. CoreWeave's release positioned the deployment for inference-intensive, agentic AI workloads, and continuous inference sessions. DatacenterDynamics and SiliconANGLE reported Michael Dell's confirmation via LinkedIn post that a liquid-cooled Dell PowerEdge XE9812 was delivered for CoreWeave, with SiliconANGLE citing theCUBE Research chief analyst's comments on broader infrastructure implications.
According to CoreWeave's press release, the Vera Rubin NVL72 configuration is fully liquid-cooled, utilizes cable-less modular trays, and completed "rigorous system-level verification" for rack-scale operations. The materials claim that rack-scale metrics include performance improvements of up to 10x per watt for inference, with reduced GPU count and cost per million tokens compared to previous generations; DatacenterDynamics reported that NVIDIA has claimed Rubin provides approximately 5x inference and 3.5x training performance improvements compared to the Blackwell generation. CoreWeave's blog and press materials also emphasized its observability and operational capabilities, including cluster-level telemetry and support engineering tailored for large-scale inference clusters.
Editorial Analysis: Public reporting positions this milestone as part of broader "New Cloud" and vendor collaborative engineering initiatives, wherein first-wave cloud providers and OEM partners validate next-generation rack-scale systems. Companies building for inference-dominant, agentic workloads increasingly prioritize liquid cooling, high-bandwidth interconnect, and integrated DPU/SuperNIC to reduce latency and per-token energy consumption. Observers cited by SiliconANGLE contend that this combination of hardware and platform engineering aims to reduce the total cost of ownership for continuous inference and large-context workloads.
Editorial Analysis: For machine learning engineers and infrastructure teams, the validated NVL72 rack represents more accessible, rack-scale inference capacity with higher per-watt token throughput. In practice, this shifts some operational focus from pure GPU count to rack-level cooling, network fabric design, and DPU offload for data movement and telemetry. Teams evaluating persistent agents or extremely long-context inference should factor rack-scale system characteristics into their benchmark planning and cost modeling.
Editorial Analysis: Observers will seek independent benchmarking beyond vendor claims, broader availability across cloud providers, and how software stacks adapt to million-token contexts and persistent sessions. Key metrics include MLPerf inference results on Vera Rubin hardware, integration of DPU/SuperNIC with orchestration and security tools, and customer case studies reporting actual token cost and latency improvements.
DatacenterDynamics republished Michael Dell's LinkedIn comment, "The world's first Nvidia Vera Rubin NVL72 server rack has arrived," with attribution to Dell. CoreWeave's press release includes a customer quote from Jane Street's Craig Falls, Head of Quantitative Research, describing performance and support advantages when scaling across previous NVIDIA generations.
Editorial Analysis: Vendor materials showcase performance and cost data; independent validation and third-party benchmarking remain necessary to quantify actual benefits for specific workloads.