Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Google launches commercial Cloud TPU service for general availability, bringing custom AI accelerators into cloud market.

Hyperscaler entry with proprietary TPU chips validates custom silicon strategy and directly competes with NVIDIA GPU dominance.
Trade pressSlicast · February 13, 2018 · Global · Source: forbes.com
importance 85

Google recently announced that the Google Cloud TPU, first announced the prior May, is now available in limited quantities as a beta on the Google Cloud Platform (GCP) for running TensorFlow-based AI applications. The search giant claims the device will substantially reduce its datacenter footprint and cost outlays. However, the company's 4-chip beast, replete with 64 GB of expensive High Bandwidth Memory (HBM), is roughly 33% more expensive per unit of performance than NVIDIA's single chip, 1-year old Tesla V100 GPU accelerator.

Google launched its TensorFlow Processing Unit one week after NVIDIA's 2017 GTC event, where the graphics chipmaker unveiled its mammoth Volta GPU. The concept of an ASIC (Application Specific Integrated Circuit) for Machine Learning has clear appeal. From a raw performance standpoint, chip to chip, the Cloud TPU chip itself delivers 45 Trillion Operations Per Second (TOPS)—more than twice the performance of NVIDIA's Pascal GPU Accelerator. However, NVIDIA's Volta GPU delivers 125 TOPS, making it almost 3 times faster (125 vs. 45 TOPS), thanks in part to its TensorCore feature that performs a 4x4 matrix multiply in a single clock cycle—provided models can take advantage of this capability.

The pricing comparison is stark: the 4-die Google Cloud TPU costs over twice as much as the NVIDIA Volta GPU available on the Amazon AWS cloud while delivering ~67% more performance, based on training time for the ResNet50 neural network. Net result: the Google part costs ~33% more to do the same work. A 4-die Cloud TPU can get training done faster, but it is >2X slower than a 4 GPU instance on AWS. The high costs are believed to be driven partly by HBM2 memory chips, estimated at over $300 per TPU die—or $900 more than the 16GB needed for Volta.

Google also announced that its TPU Pods, interconnected TPUs forming a massive compute cluster, would be ready later that year. The company had already announced that a 1000 TPU Pod, called the TensorFlow Research Cloud, would be available at no charge to the "world's top researchers" to help drive innovation on TensorFlow-based AI models. The TPU silicon only supports TensorFlow, Google's open sourced Machine Learning Framework, whereas NVIDIA's GPUs support virtually every machine learning software package. Google's TPU hardware is only available from Google as a service.

Nine months since the initial announcement, the TPU is just now becoming available as a beta service in "limited quantities." The real disappointment comes from the pricing strategy: why would anyone pay more to get the same job done? Even accounting for the HBM2 memory delta, Google is getting the ASIC at manufacturing costs, so the higher pricing cannot be fully explained by component costs alone. As advanced silicon is hard to do—even if you are Google—hopefully these prices will come down once Google irons out whatever wrinkles are limiting the quantities.

Read the original
Google launches commercial Cloud TPU service… · Slicast