Google Cloud commits to doubling down on Nvidia GPU offerings for machine learning inference workloads.
Since Google announced its own chip to accelerate deep learning, the TensorFlow Processing Unit (TPU), many industry observers questioned whether such custom chips could significantly diminish NVIDIA's influence. However, Google recently expanded its offering of NVIDIA's latest GPU, the Turing-based T4, for global availability on the Google Cloud Platform, signaling that Google's customers are expressing a preference for a cost-effective GPU for training and inference. This move suggests Google recognizes both the demand for specialized GPU acceleration alongside its own proprietary chip technology.
The market for AI inference processing—in which trained neural networks are used to "infer" or predict properties of new input data—has garnered tremendous attention as the next major opportunity for specialized semiconductors. NVIDIA estimates that 80-90% of the cost of neural networks lies in inference processing rather than training. Inference, while still a significant computational task, requires orders of magnitude less processing than training and can even be handled with far more efficient integer math, using 8 or fewer bits of precision. This shift has profound implications as the inference market, which covers applications from data centers to drones, is expected to overtake training revenue over the next decade.
NVIDIA currently dominates AI acceleration with approximately $3 billion in annual revenue, primarily from training deep neural networks—a monstrously massive processing task requiring trillions of floating-point calculations. Most inference processing today is executed on Intel Xeon processors in public clouds, though dozens of startups are building chips for more complex inference workloads. Intel has acknowledged inference acceleration as an emerging trend and added a Nervana-based inference chip to its roadmap. This competitive landscape raises critical questions about whether the industry needs accelerators beyond fast CPUs for inference and whether general-purpose GPUs like NVIDIA can compete against custom application-specific processors such as Google's TPU or Intel's Nervana.
Google's expansion of T4 support across all regions of the Google Cloud Platform demonstrates significant global demand for inference GPUs. The T4 is a versatile product supporting all AI frameworks, deep learning models, and machine learning algorithms for both training and inference. Priced as low as $0.29 per hour per GPU on GCP, the T4 costs over 75% less than an NVIDIA V100, which is priced at $1.24 for preemptable access. NVIDIA continues to expand from data center training into the inference market, pointing to services such as Snap's monetization algorithm and Microsoft Bing's conversational and image search services—all running on NVIDIA GPUs—to demonstrate GPU advantages for inference processing.
Despite inventing the TPU, Google demonstrates market awareness by serving what its customers want: fast, affordable inference processing. As AI becomes more pervasive and applications increasingly combine multiple neural networks to provide intuitive user interfaces, these services will likely run on accelerators. NVIDIA still has significant work ahead to fully demonstrate GPU advantages in inference processing within the data center, yet even Google, the inventor of the TPU, recognizes the value and market demand for inference GPUs.