엔비디아 베라 루빈 NVL72 시스템은 MLPerf Inference v6.1 벤치마크 데뷔에서 선도적인 성능을 달성하며 효율적인 인프라 확장 및 지속적인 소프트웨어 최적화를 입증했다.
System performance, efficient infrastructure scaling, and continuous software optimization are the key levers that determine AI inference economics. Higher system performance yields more generated tokens, driving greater revenue. Efficient scaling ensures throughput grows proportionally with added hardware, minimizing resource requirements for serving users at scale. Continuous optimization maximizes the return on infrastructure investments.
Underpinning these three factors is platform fungibility: a single infrastructure can run any model or workload—from training to inference, recommenders to reasoning, and language to video—keeping utilization consistently high.
The NVIDIA platform is purpose-built to optimize across all three, as demonstrated by the MLPerf Inference v6.1 results released today:
**NVIDIA Vera Rubin NVL72 Debuts With Leading Performance:** In its first MLPerf Inference preview submission, the Vera Rubin NVL72 delivers up to 3.7x higher throughput than the GB300 NVL72.
**NVIDIA GB300 NVL72 Scales With Leading Efficiency:** A 288-GPU submission spanning four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline.
**Continuous Software Optimizations Drive Performance Gains:** Software enhancements in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance compared to v6.0. These optimizations extended beyond the v6.1 submission deadline, yielding additional performance gains.
For organizations evaluating AI infrastructure, performance, scaling efficiency, and software velocity are critical determinants of long-term inference economics.
**Vera Rubin NVL72 Makes MLPerf Inference Debut With Leading Performance**
NVIDIA submitted Vera Rubin NVL72 preview results on two of the most demanding benchmarks in the MLPerf Inference v6.1 suite: DeepSeek-R1 and Qwen3-VL.
On the Qwen3-VL benchmark, Vera Rubin NVL72 achieves up to 3.7x higher throughput than GB300 NVL72 across offline, server, and interactive scenarios, leveraging vLLM with the NVIDIA Dynamo open-source inference framework. On DeepSeek-R1, which utilizes the NVIDIA TensorRT-LLM library, throughput reaches up to 2.5x higher than GB300 NVL72. These preliminary results highlight NVIDIA’s accelerated innovation cycle and demonstrate how ongoing software optimizations will further elevate performance.
MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0106 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.
This translates to tangible business outcomes: each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users, and generates greater revenue than a GB300 NVL72 rack, all while reducing cost per token.
These results reflect full-stack co-design across hardware and software. Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference. Additionally, NVFP4 precision reduces the memory footprint across model weights, attention mechanisms, and KV caches, boosting throughput with minimal impact on output quality.
The Vera Rubin submissions heavily leveraged disaggregated serving, decoupling prefill and decode workloads alongside large-scale expert parallelism. This approach maximizes efficiency across the mixture-of-experts layers powering models such as DeepSeek-R1 and Qwen3-VL.
The NVL72 scale-up domain provides the interconnect foundation necessary to make these techniques effective at rack scale. Powered by sixth-generation NVIDIA NVLink and NVLink Switch, it delivers 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet.
This co-design philosophy extends to NVIDIA’s partner ecosystem; Nebius also submitted Vera Rubin NVL72 preview results, demonstrating strong performance.
AI agents—which reason, plan, and execute across multiple steps—are reshaping how inference performance is measured. In benchmarks designed to capture this evolution, such as SemiAnalysis AgentX, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 during preview testing. Furthermore, the upcoming MLPerf Endpoints benchmark will introduce standardized measurement for agentic inference workloads, extending evaluation beyond traditional throughput metrics.
**NVIDIA GB300 NVL72 Scales With Leading Efficiency**
Scaling efficiency—how effectively additional GPUs translate into throughput gains—is a critical metric for AI infrastructure productivity. NVIDIA achieves this through high-bandwidth, low-latency scale-up interconnects within each rack, high-bandwidth networking between racks, and efficient request orchestration across nodes.
In its DeepSeek-R1 (DSR1) submission, NVIDIA scaled from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario. Throughput increased nearly in direct proportion to the added hardware.
MLPerf Inference v6.1, Closed Division. Results retrieved from www.mlcommons.org on Sep 16, 2026. NVIDIA platform results from the following entries: 6.1-0073 and 6.1-0074. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.
Scaling efficiency is critical because adding GPUs does not automatically yield proportional throughput gains. If doubling the GPU count resulted in only a single-digit percentage improvement, infrastructure costs would vastly outpace performance returns. Therefore, architecture, interconnects, and software must scale in unison.
GB300 NVL72 also demonstrated rack-scale efficiency on the WAN 2.2 text-to-video benchmark, generating 0.65 720p videos per second at 5.7 seconds per video. This represents 9x higher throughput and 7.5x lower latency compared to a single node.
**Software Optimizations Drive Continuous Gains**
The NVIDIA platform undergoes continuous software development, consistently delivering performance and feature enhancements.
In v6.1, GB300 NVL72 performance on Qwen3-VL improved by up to 1.6x compared to v6.0 results. These gains were driven by lower KV cache precision, expanded kernel fusion, optimized kernels, and disaggregated serving via vLLM and NVIDIA Dynamo.
Software optimization efforts extended beyond the v6.1 submission deadline. Unverified post-submission results for GPT-OSS-120B and DLRMv3 indicate additional performance gains.
**AI Inference at Every Scale**
Beyond the Grace Blackwell and Vera Rubin NVL72 platform results, NVIDIA also submitted Jetson AGX Thor performance data. Using NVIDIA TensorRT Edge-LLM on the newly introduced Edge-Agentic benchmark with Qwen3.6-27B, the edge device demonstrated strong capabilities.
The NVIDIA partner ecosystem participated broadly, with 19 partners—eight operating on multi-node Blackwell NVL72 systems—demonstrating strong performance. Participating partners include ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, and Wiwynn.
From compact edge devices to the largest AI factories, NVIDIA continues to advance performance across the entire technology stack. This progress is driven by an annual cadence of new platform architectures, continuously evolving software, and an ecosystem engineered to deliver solutions at scale.
Learn more about the NVIDIA Vera Rubin platform.