Friday, September 18, 2026
AI Infrastructure · News & Analysis
HomeHeadlinesReport
Headlines · Report

NVIDIA Vera Rubin NVL72 systems achieved up to 3.7 times better inference throughput than NVIDIA GB300 NVL72 in their debut submission to MLPerf Inference v6.1.

The performance leap validates NVIDIA’s next-generation silicon roadmap and sets a new baseline for hyperscaler inference cluster procurement specs.
Trade pressSlicast · September 17, 2026 at 13:17 UTC · Global · Source: HPCwire
importance 85

Sept. 17, 2026 — System performance, efficient infrastructure scaling, and continuous software optimization are the primary levers that determine AI inference economics. Higher system performance yields more generated tokens, which directly translates to higher revenue. Efficient scaling ensures that throughput grows proportionally as hardware is added, minimizing the resources required to serve users at scale. Continuous optimization maximizes the value extracted from infrastructure investments. Underpinning all three is platform fungibility: the same infrastructure can run any model and workload—spanning training to inference, recommender systems to reasoning, and language to video—keeping utilization high.

The NVIDIA platform is purpose-built to optimize across these dimensions, as demonstrated by the MLPerf Inference v6.1 results released today. In its first MLPerf Inference preview submission, the NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than the GB300 NVL72. Meanwhile, a 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline. Additionally, software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0, with further gains realized post-submission. For organizations making AI infrastructure decisions, performance, scaling efficiency, and software velocity remain critical considerations that shape long-term inference economics.

NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL. On Qwen3-VL, using vLLM with the NVIDIA Dynamo open-source inference framework, Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 across offline, server, and interactive scenarios. On DeepSeek-R1, utilizing the NVIDIA TensorRT-LLM library, throughput is up to 2.5x higher than GB300 NVL72. These early results highlight NVIDIA’s accelerated pace of innovation and demonstrate how performance will continue to improve through ongoing software optimization. Consequently, each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users, and generates more revenue than a GB300 NVL72 rack while lowering cost per token.

This performance reflects full-stack co-design across hardware and software. Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference, while NVFP4 precision reduces the memory footprint across model weights, attention mechanisms, and KV caches, increasing throughput with minimal loss of output quality. The Vera Rubin submissions heavily leveraged disaggregated serving, separating prefill and decode alongside large-scale expert parallelism to maximize efficiency across the mixture-of-experts layers powering models like DeepSeek-R1 and Qwen3-VL. The NVL72 scale-up domain, powered by sixth-generation NVIDIA NVLink and NVLink Switch, delivers 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet, providing the interconnect foundation that makes these techniques effective at rack scale. This co-design extends to NVIDIA’s partner ecosystem; Nebius also submitted Vera Rubin NVL72 preview results and demonstrated excellent performance.

AI agents, which reason, plan, and act across multiple steps, are reshaping how inference performance is measured. In benchmarks designed to capture this shift, such as SemiAnalysis AgentX, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing. Furthermore, the upcoming MLPerf Endpoints benchmark will introduce standardized measurement for agentic inference workloads, extending beyond what traditional throughput benchmarks capture.

Scaling efficiency—how effectively additional GPUs translate into throughput gains—is a key measure of AI infrastructure productivity. NVIDIA achieves this through high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks, and efficient request orchestration across nodes. NVIDIA’s DeepSeek-R1 (DSR1) submission scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario. Throughput grew nearly in proportion to the hardware added. Scaling efficiency is critical because adding more GPUs does not automatically yield proportional throughput gains. If doubling the GPU count resulted in only a single-digit percentage improvement, infrastructure costs would far outpace performance returns. The architecture, interconnect, and software must scale together. GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, reaching 0.65 720p videos per second at 5.7 seconds per video, representing 9x higher throughput and 7.5x lower latency than a single node.

The NVIDIA platform undergoes continuous software development, delivering ongoing performance and feature improvements. In v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 results. These gains were achieved through lower KV cache precision, additional kernel fusion, optimized kernels, and disaggregated serving with vLLM and NVIDIA Dynamo. Software optimization continued past the v6.1 submission deadline as well. Post-submission results, which have not yet been verified by MLCommons, on GPT-OSS-120B and DLRMv3 show further performance gains.

Beyond the NVIDIA Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA submitted Jetson AGX Thor results using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B. The NVIDIA partner ecosystem participated broadly, with 19 partners—eight of them on multi-node Blackwell NVL72 systems—demonstrating excellent performance. This includes ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, and Wiwynn. From compact edge devices to the largest AI factories, NVIDIA continues to advance performance across the full technology stack with an annual cadence of platform architectures, continuously improving software, and an ecosystem built to deliver it at scale.

Read the original
NVIDIA Vera Rubin NVL72 systems achieved up to… · Slicast