Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeCompute & CloudReport
Compute & Cloud · Report

NVIDIA is strategically prioritizing AI inference workloads and aggressive expansion into the Chinese market as key growth frontiers.

Inference represents higher-volume, lower-cost AI workload than training; combined with China dominance, reshapes NVIDIA's product roadmap and market strategy.
Trade pressSlicast · September 28, 2017 · Global · Source: forbes.com
importance 76

NVIDIA's datacenter business, now generating approximately $1.6B annually, has been driven primarily by demand to train deep neural networks for Machine Learning and Artificial Intelligence applications. Much of this business originates from the largest U.S. datacenters operated by Amazon, Google, Facebook, IBM, and Microsoft. At its annual Beijing Graphics Technology Conference (GTC) event, NVIDIA announced new technology and customer initiatives aimed at capturing market share in the inference segment and strengthening its position in China's AI market. Inference—where trained neural networks predict and classify sample data—is expected to eventually exceed the training market in terms of chip unit volumes, as organizations transition from research and development to commercial deployment both in cloud and edge environments.

NVIDIA CEO Jensen Huang announced TensorRT3 software designed to optimize trained neural networks for inference processing on NVIDIA GPUs. TensorRT3 packages neural networks built with any ML framework for deployment across NVIDIA's datacenter and edge device portfolio, functioning as "essentially the CUDA of inferencing." The software is now being deployed by China's largest Internet datacenters—Alibaba, Baidu, Tencent, and JD.com—for ML workloads. Huang presented compelling performance benchmarks demonstrating the company's capabilities, highlighting approximately a 20X increase in performance for inference processing of images using ResNet-50 with the new NVIDIA software compared to previous V100 (Volta) results.

Complementing the TensorRT3 announcements, Alibaba, Baidu, and Tencent—the largest Chinese Cloud Service Providers—announced they are offering NVIDIA's newest Tesla V100 GPUs to customers for scientific and deep learning applications. For enterprises deploying deep learning in their own datacenters, Huawei, Inspur, and Lenovo announced they would sell HGX-based servers with Volta GPUs to their global customer base. The HGX platform, an 8-GPU chassis with NVLink interconnect designed to provide high GPU scaling density, was designed with Microsoft and is available as an open source hardware platform through the Open Compute program. Lenovo, now the only global OEM offering HGX, represents a particularly significant win as the company seeks high-density GPU servers for large-scale training workloads.

Beyond datacenter applications, NVIDIA announced that JD.com's delivery subsidiary would deploy the NVIDIA Jetson platform to guide and control land and air drone delivery services. As a solution to challenges posed by China's congested urban transportation infrastructure, JD.com plans to deploy approximately one million Jetson-equipped drones in service by 2020. As Machine Learning transitions from research to commercial deployment, NVIDIA is positioning itself to compete with CPUs, FPGAs, and ASICs across varying datacenter and edge processing requirements. The announced customer wins demonstrate NVIDIA's competitive position; however, unlike training—where NVIDIA maintains dominance—the diversity of inference data, latency, and power requirements is expected to create a competitive landscape with multiple viable solutions.

Read the original
NVIDIA is strategically prioritizing AI… · Slicast