MLPerf inference benchmarks confirm Nvidia performance leadership; Jetson edge product line expands market reach.
The MLPerf consortium has released the first wave of results from its Inference v0.5 benchmark, which consists of five massively-parallel benchmarks focused on three common machine learning tasks including Image Classification, Object Detection, and Machine Translation. These benchmarks are targeted at applications ranging from autonomous driving and facial recognition to natural language processing, scaling across form factors from smartphones and PCs to cloud computing platforms in the data center. NVIDIA, Google, Intel, Alibaba and others from around the globe submitted results, with NVIDIA submitting the broadest range of results spanning the widest array of scenarios. NVIDIA's Turing-based GPUs and Jetson Xavier platform performed comparatively well.
In the ImageNet ResNet-50 v1.5 Offline scenario, only a limited number of direct comparisons can be made since many companies submitted only a few results. NVIDIA demonstrated strong per-processor performance with its Titan RTX, scoring 65,431.40 inputs/second. Many submissions are from systems with multiple processors; the SCAN 3XS DBP T496X2 Fluid system uses four Titan RTX cards and scored 66,250.40 inputs/second, while the 2x Google Cloud TPU v3-8, which leverages eight processors, achieved similar results. NVIDIA's performance drops by only a few percentage points when moving from the Offline to Server scenario for the ImageNet ResNet-50 v1.5 test, whereas Google's performance is slashed nearly in half in the same transition. Note that the Titan RTX doesn't support ECC memory, so while its performance is strong, its lack of this feature may preclude its use by some data center customers. Habana Labs' HL-102-Goya PCI-board also performed well with a single-processor result of 14,451.00 inputs/second in the ImageNet ResNet-50 v1.5 Offline scenario. A pair of Intel Nervana NNP-I 1000s scored 10,567.20 in the same test and showed the smallest performance decrease moving from the Offline to Server scenario (10,262.63 QPS vs 10,567.20 inputs/second).
Ian Buck, general manager and vice president of Accelerated Computing at NVIDIA, stated that "AI is at a tipping point as it moves swiftly from research to large-scale deployment for real applications" and that "AI inference is a tremendous computational challenge. Combining the industry's most advanced programmable accelerator, the CUDA-X suite of AI algorithms and our deep expertise in AI computing, NVIDIA can help data centers deploy their large and growing body of complex AI models." The results do not account for pricing or power consumption, and different form factors are employed—add-in accelerators, servers, the cloud, and others—though the data does level set expectations and highlight some standout performers.
In addition to its strong MLPerf results, NVIDIA announced a new Jetson Xavier NX device for AI computing on the edge. The new device essentially shrinks Jetson Xavier-class performance down into a much smaller form factor built around a low-power version of the Xavier SoC. Targeting robotic and embedded computing devices at the edge, the Jetson Xavier NX features a compact 70mm x 45mm form factor. The SoC powering the module includes a 6-core Carmel Arm 64-bit CPU with 6MB of L2 and 4MB of L3 cache, a Volta-based GPU with 384 CUDA cores, 48 Tensor cores, and two NVIDIA Deep Learning Accelerators. It features 8GB of LPDDR4x memory connected via a 128-bit interface offering up to 51.2GB/second of peak bandwidth. The Jetson Xavier NX supports up to six CSI cameras (36 via virtual channels), Gigabit Ethernet, and can encode 2 x 4K30 video streams and decode 2 x 4K60 streams.