Thursday, September 17, 2026
AI 인프라 · 뉴스 & 분석
반도체·하드웨어리포트
반도체·하드웨어 · 리포트

MLCommons has released MLPerf Inference v6.1 benchmark results, which introduce two new tests for emerging AI deployment patterns including Agentic Inference.

The expanded benchmark suite provides standardized performance validation for next-generation inference workloads, guiding accelerator procurement strategies.
업계 전문지Slicast · 2026년 9월 16일 17:21 UTC · 글로벌 · 출처: HPCwire
중요도 65

SAN FRANCISCO, Sept. 16, 2026 — MLCommons announced new results for its industry-standard MLPerf Inference v6.1 benchmark suite. Setting a new record for participating organizations, this release introduces two new tests aligned with emerging AI inference deployment trends. It also delivers the first peer-reviewed performance results for several recently launched or imminent AI platforms, demonstrating up to a 5.7X performance gain compared to results from one year ago.

The open-source MLPerf Inference benchmark suite evaluates system performance in an architecture-neutral, representative, and reproducible manner. By establishing a standardized testing environment, it drives innovation, performance, and energy efficiency across the industry. These published results deliver critical technical insights, enabling customers procuring and deploying AI systems to make informed decisions grounded in trusted empirical data.

MLPerf Inference v6.1 expands the suite with two new tests that reflect the industry’s shift toward more complex, multi-step, and agentic AI inference deployments across both datacenter and edge environments. The End-to-End Retrieval-Augmented Generation (RAG) benchmark evaluates a complex pipeline of multiple distinct AI models for question answering. This workload decomposes the task into sequential steps: an embedding model converts a query into a vector, a retriever pulls candidate passages from a vector database, a re-ranker refines the list, and one or more large language models reason over the data to generate an answer. The test captures performance data for two related tasks: ingesting a document corpus to build a vector database, and executing query-answering against a pre-built database. Additional details on the End-to-End RAG benchmark are available online.

The Edge Agentic Inference benchmark captures the industry’s move from single-shot interactions to complex, multi-turn workloads such as agentic coding. Unlike single-shot inference, agentic workloads maintain a growing conversational history where each query builds upon previous exchanges. The model continuously merges evidence gathering with iterative reasoning to reach a final result. This usage model unlocks significant new capabilities while substantially increasing compute demands. As agentic AI deployment grows at the edge, operators face unique performance challenges when serving individual users with fixed memory, processing, power, and context limits. Adopting the framework and methodology from the forthcoming MLPerf Agentic datacenter benchmark, this test specializes the approach for edge environments by incorporating an edge model and quantization, a single-stream coding workload, latency metrics, and a statistically robust accuracy gate. It comprises two tasks: measuring workload accuracy under strict time constraints, and evaluating the performance of a deterministic workload where accuracy checks are intrinsic to the process. Further information on the Edge Agentic Inference benchmark is available online.

MLPerf Inference v6.1 also integrates support for speculative decoding, a widely adopted production optimization that predicts and verifies multiple tokens within a single forward pass. This capability is now supported in the interactive scenario for two inference benchmarks, with additional implementation in the GPT-OSS task.

“We are working hard to ensure that the MLPerf Inference benchmark continues to reflect the scenarios that the AI community values most,” said Miro Hodak, MLPerf Inference working group co-chair. “We added the End-to-end RAG test because it’s clear that query-answering has evolved beyond simply an LLM trained on a corpus; stakeholders need to understand the real-world performance of the types of multi-step, multi-component pipelines that are being built today. Likewise, we added the Edge Agentic Inference test because complex inference systems with agentic properties are increasingly hosted on edge computing devices, creating a new set of performance challenges our customers face today. We are committed to providing timely and relevant performance data that reflects real-world production systems and the performance optimizations that are being deployed today.”

This submission cycle underscores the rapid evolution of AI technology, with participants showcasing a diverse array of hardware configurations spanning accelerators, boards, workstations, and servers. Five new processors and accelerators were evaluated: the commercially available AMD Ryzen AI Max+ 395, AMD Instinct MI350P, and Intel Arc Pro B70, alongside the NVIDIA Rubin and NVIDIA Vera Rubin NVL72, which are currently in preview. The submissions also feature the largest system ever tested in the MLPerf Inference benchmark, comprising 512 accelerators and two novel heterogeneous architectures. The first pairs high-performance networking with accelerators from competing vendors, while the second spans the Pacific Ocean in a geographically distributed configuration.

Performance gains continue to compound, repeatedly breaking records for AI accelerator capabilities. In the Visual Language Model (VLM) test, the top per-accelerator server scenario improved by 2.99x compared to v6.0 results from just six months prior. Similarly, the Deepseek R1 test recorded a 5.7x improvement over v5.1 results from one year ago. These compounding performance gains directly translate to expanded capabilities for end customers, enabling higher user concurrency and more sophisticated inference workloads.

“With the critical data from the MLPerf Inference benchmark, the AI community is once again proving that what can be measured can be improved,” said Frank Han, MLPerf Inference working group co-chair. “As workloads and the underlying systems increase in capability, scale, and complexity across a broader variety of form factors and physical configurations, we are clearly seeing a renewed interest in ensuring that inference can take full advantage of hardware, software, and architectural advances – both in the datacenter and at the edge. With performance data from the Inference v6.1 benchmark, customers can better understand the cost-benefit tradeoffs and make informed decisions on how to procure and deploy their AI systems.”

The MLPerf Inference v6.1 benchmark attracted a record 30 participating organizations: AMD, ASUSTeK, Atlas Inference, Cisco, CoreWeave, Crusoe, Dell, Fujitsu Limited, GigaComputing, Google, Hewlett Packard Enterprise, Intel, Inventec Corporation, KRAI, Lambda, MangoBoost, Microsoft Azure, MiTAC, Nebius, NVIDIA, Oracle, Orrick, Quanta Cloud Technology, RedHat, ScitiX, Supermicro, Telecommunications Technology Association, VibeHPC, Wiwynn, and individual contributor Naeem Khoshnevis.

“I would like to welcome our six first-time submitters: Atlas Inference, Crusoe, Orrick Industries LLC, ScitiX, VibeHPC, and individual contributor Naeem Khoshnevis,” said David Kanter, Head of MLPerf. “The groundswell of participation from the AI community tells us that our work is important and is making a difference to stakeholders.”

Additionally, more than half of all submitters utilized MLPerf’s new API-centric harness, which serves as the foundation for the upcoming MLPerf Endpoints benchmark suite. This harness employs a true client-server architecture over industry-standard APIs to transmit inference queries and results, offering a more accurate reflection of real-world datacenter deployments. “We’re excited to see submitters embrace our customer-centric tooling and workloads and partner with us to improve the state of industry-wide benchmarking for AI,” said Kanter. “Moving forward, MLPerf Endpoints will replace Inference in our family of benchmarks for the datacenter, and the quick uptake of our API-centric harness will contribute to making that transition seamless.”

To review the MLPerf Inference v6.1 results, visit the official Datacenter and Edge benchmark pages and consult the supplemental documentation for detailed submission breakdowns. An interactive visualization of the datacenter results is available via the MLPerf benchmark dashboard at https://mlcommons.org/visualizer. Submitters have also provided supplementary statements to contextualize their methodologies and findings.

MLCommons is the global leader in AI benchmarking. Supported by over 130 member organizations and affiliates, this open engineering consortium consistently unites academia, industry, and civil society to measure and advance artificial intelligence. Established in 2018 around the initial MLPerf benchmarks, the organization quickly expanded into a comprehensive suite of industry standards designed to quantify machine learning performance and promote technical transparency. Through collaborative engineering, MLCommons continues to develop the benchmarks and metrics necessary to evaluate AI systems for accuracy, safety, speed, and efficiency, empowering customers to navigate trade-offs and make informed procurement decisions.

원문 보기
MLCommons has released MLPerf Inference v6.1… · Slicast