Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Google disclosed technical specifications of TPU v4 datacenter systems currently in production.

Hyperscaler custom silicon strategy confirmed with open disclosure; validates multiple-path accelerator competition and shapes AI chip landscape.
Trade pressSlicast · April 6, 2023 · Global · Source: theregister.com
importance 80

Google on Wednesday revealed additional details about its fourth-generation Tensor Processing Unit chip (TPU v4), claiming superiority over Nvidia's established hardware. According to researchers from Google and UC Berkeley, TPU v4 "is 1.2x–1.7x faster and uses 1.3x–1.9x less power than the Nvidia A100," as stated in a paper published ahead of a June presentation at the International Symposium on Computer Architecture. Nvidia responded with its own claims, with founder and CEO Jensen Huang noting that the A100 debuted three years ago and that the company's more recent H100 (Hopper) GPUs deliver 4x more performance than A100 based on MLPerf 3.0 benchmarks.

The comparison reflects a timing decision by Google researchers. Both TPU v4 and A100 were deployed in 2020 and use 7nm technology. The Google and UC Berkeley authors chose not to measure TPU v4 against the H100, explaining: "The newer, 700W H100 was not available at AWS, Azure, or Google Cloud in 2022. The appropriate H100 match would be a successor to TPU v4 deployed in a similar time frame and technology (e.g., in 2023 and 4nm)." The TPU v4 represents the company's fifth domain specific architecture tuned for machine learning and its third supercomputer for ML models, and it outperforms its v3 predecessor by 2.1x while boasting 2.7x better performance per Watt.

The key innovations in TPU v4 center on two technologies: Optical Circuit Switches (OCS) with optical data links and SparseCores (SC), dataflow processors that accelerate calculations for embedding-dependent models like recommender systems. OCS interconnection hardware allows Google's 4K TPU node supercomputer to operate with 1,000 CPU hosts that are occasionally unavailable (0.1–1.0 percent of the time) without problems, with researchers noting that "an OCS raises availability by routing around failures." Host availability must reach 99.9 percent without OCS, but with OCS, effective throughput in the TPU supercomputer can be achieved with availability around 99.0 percent. SC processors "accelerate models that rely on embeddings by 5x–7x yet use only five percent of die area and power," according to the researchers, and are particularly valuable given that embedding-dependent deep learning recommendation models represent a quarter of Google's workloads in advertising, search ranking, YouTube, and Google Play applications.

When 4,096 TPU v4 nodes are unified into a supercomputer in a datacenter, the resulting hardware requires approximately 2–6x less energy and approximately 20x less carbon dioxide emissions than rival domain-specific architectures, the researchers claim. The authors declare that "a ~20x reduction in carbon footprint greatly increases the chances of delivering on the amazing potential of ML in a sustainable manner." Google has dozens of these supercomputers deployed for both internal and external use across its cloud services.

Read the original
Google disclosed technical specifications of… · Slicast