Thursday, September 17, 2026
AI 인프라 · 뉴스 & 분석
반도체·하드웨어리포트
반도체·하드웨어 · 리포트

NVIDIA Vera Rubin NVL72가 첫 MLPerf 추론 프리뷰 결과를 공개했다.

초기 벤치마크 공개를 통해 추론 처리량에서 아키텍처 개선이 확인되었으며, 이는 차세대 엔터프라이즈 AI 학습 및 서빙 워크로드에 대한 준비가 완료되었음을 입증한다.
업계 전문지Slicast · 2026년 9월 16일 15:24 UTC · 미국 · 출처: Unite.AI
중요도 85

On September 16, 2026, NVIDIA published its MLPerf Inference v6.1 submission results, headlined by the first preview submission for its Vera Rubin NVL72 system. The company reported that the platform delivers up to 3.7 times higher throughput than the GB300 NVL72 system on the Qwen3-VL benchmark.

MLPerf Inference: Datacenter is an MLCommons Association benchmark suite designed to measure how quickly systems process inputs and generate outputs using trained models. According to MLCommons’ published definitions, the Closed division—the category NVIDIA used for its headline figures—requires participants to use the same model as the reference implementation, enabling apples-to-apples comparisons across hardware platforms and software frameworks. Systems listed in the Preview availability category must be commercially available by the next submission round.

NVIDIA submitted preview results for Vera Rubin NVL72 on two of the most demanding benchmarks in the v6.1 suite: DeepSeek-R1 and Qwen3-VL. On Qwen3-VL, the platform achieved up to 3.7 times higher throughput than the GB300 NVL72 across offline, server, and interactive scenarios, leveraging vLLM alongside NVIDIA’s open-source Dynamo inference framework. For DeepSeek-R1, which utilized the NVIDIA TensorRT-LLM library, Vera Rubin NVL72 delivered up to 2.5 times higher throughput compared to the GB300 NVL72.

The Closed division figures, sourced from mlcommons.org on September 16, 2026, correspond to submission entries 6.1-0106 and 6.1-0074. NVIDIA stated that these performance gains enable each Vera Rubin NVL72 rack to deliver significantly more tokens, serve more users, and generate greater revenue than a GB300 NVL72 rack, all while reducing cost per token. Nebius also submitted Vera Rubin NVL72 preview results in the same round, which NVIDIA characterized as demonstrating excellent performance.

NVIDIA attributed these results to full-stack hardware and software codesign. The company noted that Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference. Additionally, NVFP4 precision reduces the memory footprint of model weights, attention mechanisms, and KV caches, boosting throughput with what NVIDIA described as minimal impact on output quality.

The submissions leveraged disaggregated serving, which decouples the prefill and decode phases, alongside large-scale expert parallelism across the mixture-of-experts layers powering models like DeepSeek-R1 and Qwen3-VL. NVIDIA highlighted that the NVL72 scale-up domain, built on sixth-generation NVLink and NVLink Switch technology, delivers 10 times higher packet rates and 3 times lower latency than standard Ethernet, providing the necessary interconnect foundation for these techniques at rack scale.

NVIDIA also reported that Vera Rubin NVL72 outperformed the GB300 NVL72 by 30 times in preview testing on the SemiAnalysis AgentX benchmark, which is specifically designed to evaluate agentic workloads. The company added that the upcoming MLPerf Endpoints benchmark will introduce standardized measurement for agentic inference tasks, capturing capabilities that traditional throughput benchmarks cannot.

In the same submission round, NVIDIA scaled its DeepSeek-R1 workload from a single GB300 NVL72 rack of 72 GPUs to four racks totaling 288 GPUs. This configuration achieved 99 percent scaling efficiency in the offline scenario, with throughput increasing nearly in direct proportion to the added hardware (submission entries 6.1-0073 and 6.1-0074). On the WAN 2.2 text-to-video benchmark, the GB300 NVL72 generated 0.65 720p videos per second at a rate of 5.7 seconds per video. NVIDIA noted this represents 9 times higher throughput and 7.5 times lower latency compared to a single node.

NVIDIA also reported that GB300 NVL72 performance on Qwen3-VL improved by up to 1.6 times in v6.1 compared to v6.0, attributing the gains to reduced KV cache precision, additional kernel fusion, optimized kernels, and disaggregated serving via vLLM and NVIDIA Dynamo. Optimization efforts continued past the v6.1 submission deadline, yielding further gains on GPT-OSS-120B and DLRMv3. These post-deadline results remain unverified by MLCommons.

Beyond rack-scale platforms, NVIDIA submitted Jetson AGX Thor results using the TensorRT Edge-LLM library on the newly introduced Edge-Agentic benchmark with the Qwen3.6-27B model. Nineteen partners participated in the round, including eight that submitted results on multi-node Blackwell NVL72 systems: ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, and Wiwynn.

MLCommons notes on its benchmark page that published results may be modified or invalidated following initial publication, with all adjustments documented in its official change log.

Theo Nash is an AI-generated specialist at Unite.AI, covering AI infrastructure, compute, and the hardware systems that power modern artificial intelligence. His reporting focuses on the technical foundations behind large-scale AI workloads, including data centers, accelerators, networking, and the software stacks that integrate them. With an analytical and engineering-driven perspective, Nash examines how advances in GPUs, custom silicon, memory architectures, and distributed systems enable new generations of AI models. He emphasizes performance trade-offs, energy efficiency, scalability, and the practical constraints shaping real-world AI infrastructure deployment. All articles authored by Theo Nash are AI-generated and subsequently reviewed by Unite.AI’s editorial team to ensure technical accuracy, clarity, and responsible coverage of the rapidly evolving AI compute landscape.

원문 보기
NVIDIA Vera Rubin NVL72가 첫 MLPerf 추론 프리뷰 결과를… · Slicast