Thursday, September 17, 2026
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

AMD benchmarked a 512-unit MI355X GPU cluster in MLPerf 6.1, while NVIDIA maintained strong positioning with Blackwell and its first Vera Rubin submission, alongside Intel Arc Pro and Xeon entries.

Multi-vendor benchmark participation validates competitive parity in large-scale inference clusters and signals accelerating adoption of next-gen architectures across hyperscalers.
Trade pressSlicast · September 16, 2026 at 17:15 UTC · US · Source: wccftech.com
importance 85

MLPerf Inference v6.1 has been released by MLCommons, featuring a suite of workflows aligned with current AI industry trends. As customary, AMD, NVIDIA, and Intel have submitted results from their latest hardware architectures, showcasing competitive performance across multiple benchmarks.

For the DeepSeek R1 workload, AMD submitted its highest-ever cluster configuration, utilizing 512 Instinct MI355X GPUs. This setup led performance in both offline and server use cases. NVIDIA’s GB300 and GB200 architectures were evaluated in 288-GPU configurations, with the GB300 closely matching the MI355X platform. Notably, this marks the first appearance of NVIDIA’s Vera Rubin (VR200) architecture in the benchmark suite, tested in both 72-GPU and 36-GPU configurations. The 72-GPU Vera Rubin setup delivers up to 95% higher throughput than a comparable 72-GPU GB300 configuration, while the 36-GPU variant outperforms a 72-GPU GB200 setup. These early preview results indicate significant performance gains, with additional Vera Rubin data expected as deployment expands.

For GPT-OSS 120B, AMD again deployed its 512-GPU MI355X cluster, securing top positions in both offline and server scenarios. NVIDIA’s GB300 “Grace Blackwell Ultra” configuration led the 72-GPU and 8-GPU segments. AMD’s newer MI350P architecture also demonstrated improved performance over the MI300X in 8-GPU setups. Intel made a notable entry with its Arc Pro B70, while NVIDIA’s RTX PRO series was evaluated in 8- and 4-GPU configurations.

In the Llama 2 70B benchmark, NVIDIA’s Blackwell GPUs dominated the 72-GPU segment. The GB300 also surpassed AMD’s MI355X in 8-GPU configurations, though the Instinct platform still outpaced the GB200 Blackwell solution. Workstation-class hardware continued to gain traction, with AMD’s MI350P, NVIDIA’s RTX PRO series, and Intel’s Arc Pro making strong showings, underscoring the growing relevance of non-datacenter platforms in AI inference workloads. For Llama 3.1-8B, the NVIDIA GB300 (8-GPU) configuration proved up to 17% faster than the AMD MI355X (8-GPU) setup. Additional results highlighted competitive scaling across various GPU counts and architectures, including Intel’s Xeon 6 processors in multi-socket configurations.

Following the benchmark release, each vendor shared key takeaways from their submissions:

NVIDIA stated: "Vera Rubin NVL72 system debuts with leading performance: In its first MLPerf Inference preview submission, NVIDIA Vera Rubin NVL72 delivers up to 3.7x better throughput than GB300 NVL72." They added: "GB300 NVL72 scales with leading efficiency: A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline." Regarding software, NVIDIA noted: "Continuous software optimizations drive performance gains: Software optimizations in NVIDIA’s MLPerf Inference v6.1 submissions delivered up to 1.6x higher performance over v6.0. Optimizations continued post-v6.1 submission, delivering further performance gains."

AMD highlighted: "More performance from the same AMD Instinct MI355X GPU hardware: Within one MLPerf cycle, continued AMD ROCm software optimization increased 8-GPU GPT-OSS-120B throughput, reduced Wan-2.2 latency, and raised cluster throughput with fewer GPUs." They also reported: "Leading performance across AMD Instinct MI355X and MI350P GPUs: AMD Instinct MI355X GPUs led NVIDIA B200 and B300 GPT-OSS-120B results at 8 GPUs and the NVIDIA GB200 result at 72 GPUs. In its first MLPerf round, AMD Instinct MI350P GPUs led selected NVIDIA RTX PRO 6000 Server Edition and H200 NVL results." On scale, AMD noted: "Record token throughput and production scale inference with efficient scaling: At 72 GPUs, the AMD Instinct MI355X GPU achieved 95% scale efficiency on GPT-OSS-120B. Separately, Crusoe delivered record aggregate token throughput at 512 GPUs, reaching 5.75 million Offline tokens per second on GPT-OSS-120B and 2.90 million on DeepSeek-R1."

Intel outlined: "About Intel Xeon 6 Results: Intel broadened its Xeon participation in MLPerf v6.1 from two benchmarked Xeon 6 SKUs in v6.0 to five, increasing the number of CPU inference results from 24 to 35. Intel Xeon remains the only standalone server CPU represented in MLPerf Inference submissions." Regarding their accelerator, Intel stated: "About Intel Arc Pro B70 Results: Intel Arc Pro B70 submissions expand on how software improvements can extend performance across a range of AI models. A single node configured with four Intel Arc Pro B70 GPUs provides 128GB of VRAM and supported submissions across Llama 3.1 8B, Llama 2 70B, gpt-oss-120B, Whisper and end-to-end retrieval-augmented generation (E2E-RAG). On the same four-GPU system used in v6.0, gpt-oss-120B Server performance improved 36%, while Offline performance increased 27%, reflecting continued maturity of Intel’s software stack."

MLPerf Inference v6.1 illustrates a highly competitive and rapidly evolving landscape rather than a clear-cut winner. AMD’s 512-GPU MI355X cluster set the pace at extreme scale for DeepSeek R1 and GPT-OSS 120B, while NVIDIA’s initial Vera Rubin preview signals a substantial leap: a 72-GPU VR200 system nearly doubled the throughput of a matching GB300 rack, and even a 36-GPU configuration surpassed a 72-GPU GB200. At more conventional 8- and 72-GPU scales, performance varies—GB300 frequently leads in Llama 2 70B and Llama 3.1 8B, AMD remains competitive or ahead on select GPT-OSS runs, and MI350P alongside RTX PRO and Intel Arc Pro B70 confirm that workstation hardware now competes directly in enterprise AI inference. Intel further solidified Xeon’s position as the sole standalone CPU option in the suite.

The underlying narrative, however, centers on software-driven acceleration. NVIDIA reported up to 1.6× performance gains from stack refinements since v6.0; AMD extracted additional throughput from the same MI355X silicon within one benchmark cycle; and Intel lifted Arc Pro results on identical hardware. These figures represent snapshots, not ceilings. As Vera Rubin enters production and ROCm and CUDA ecosystems continue to mature, subsequent MLPerf rounds will inevitably reflect new baselines—a trajectory that aligns perfectly with industry expectations.

Hassan Mujtaba, a Software Engineer by training and PC enthusiast by passion, serves as Wccftech’s Senior Editor for Hardware. With extensive industry experience, he specializes in deep-dive technical analysis of next-generation CPU and GPU architectures, motherboards, and cooling solutions. His coverage spans breaking technology news, hands-on reviews, and comprehensive benchmarking. Follow Wccftech on Google News for continued coverage.

Read the original
AMD benchmarked a 512-unit MI355X GPU cluster… · Slicast