NVIDIA's Blackwell platform won across all MLPerf Training 6.0 benchmarks, with the newer GB300 delivering 1.6x faster t
NVIDIA announced dominant results in MLPerf Training 6.0, the latest peer-reviewed industry benchmark for AI training performance. The NVIDIA Blackwell platform led across every category and was the only platform submitted across all seven benchmarks, delivering the fastest training time on each.
The benchmark round introduced two new mixture-of-experts pretraining workloads reflecting the growing importance of MoE architectures: DeepSeek-V3 671B and GPT-OSS-20B. Large-scale MoE training requires all-to-all communication to route tokens across GPUs to the correct expert subnetwork, a challenge NVIDIA addressed through NVLink's bandwidth advantages.
NVIDIA submitted results on both GB200 NVL72 and GB300 NVL72 rack-scale systems, where fifth-generation NVLink Switches connect 72 GPUs into a unified compute and memory pool. The GB300 NVL72 delivered up to 1.6x faster training than GB200 NVL72, driven by higher compute density with NVFP4 precision training, expanded memory capacity and increased power ceiling.
NVIDIA achieved record-breaking scale: 8,192 GPUs on GB200 NVL72 systems for DeepSeek-V3 671B, marking the largest-scale Blackwell-based submission to date, and 5,120 GPUs on Llama 3.1 405B. The company also highlighted NVFP4 training methods used to pretrain its 550-billion-parameter Nemotron 3 Ultra model.
Nineteen ecosystem partners participated, including Microsoft Azure, Google Cloud, CoreWeave, Dell Technologies, Fujitsu and others. Customer examples demonstrated real-world impact: Cohere achieved 3x faster training on GB200, Thinking Machines Lab saw 2x faster speeds on GB300, and Nebius-hosted Higgsfield reduced training time by 30 percent while serving 22 million users and generating over 6 million pieces of AI content daily. Midjourney trained its v8 image generation model on a Blackwell cluster and is scaling Blackwell Ultra for upcoming projects.