Nvidia advances AI inference and edge computing capabilities beyond traditional cloud datacenter scope.
NVIDIA today posted the fastest results on MLPerf Inference 0.5, the industry's first independent suite of AI benchmarks for inference, demonstrating the performance of NVIDIA Turing™ GPUs for data centers and NVIDIA Xavier™ system-on-a-chip for edge computing. The results represent a continuation of the company's strong performance in recent AI benchmarks, following multiple MLPerf 0.6 benchmark wins for AI training in July, where NVIDIA set eight records in training performance.
MLPerf's five inference benchmarks are applied across a range of form factors and four inference scenarios, covering established AI applications including image classification, object detection, and translation. NVIDIA topped all five benchmarks for both data center-focused scenarios (server and offline), with Turing GPUs providing the highest performance per processor among commercially available entries. Xavier provided the highest performance among commercially available edge and mobile SoCs under both edge-focused scenarios (single-stream and multi-stream).
NVIDIA distinguished itself as the only AI platform company to submit results across all five MLPerf benchmarks. The company's NVIDIA GPUs accelerate large-scale inference workloads in the world's largest cloud infrastructures, including Alibaba Cloud, AWS, Google Cloud Platform, Microsoft Azure, and Tencent. World-leading businesses and organizations, including Walmart and Procter & Gamble, are using NVIDIA's EGX edge computing platform and AI inference capabilities to run sophisticated AI workloads at the edge.
All of NVIDIA's MLPerf results were achieved using NVIDIA TensorRT™ 6, a high-performance deep learning inference software that optimizes and deploys AI applications in production from the data center to the edge. New TensorRT optimizations are also available as open source in the GitHub repository. Expanding its inference platform, NVIDIA introduced Jetson Xavier NX, described as the world's smallest, most powerful AI supercomputer for robotic and embedded computing devices at the edge, built around a low-power version of the Xavier SoC used in the MLPerf Inference 0.5 benchmarks.