Friday, August 7, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Google Cloud announces sixth-generation TPUs for AI infrastructure operations.

Major hyperscaler invests in custom silicon, signaling long-term commitment to in-house AI compute and reducing dependency on Nvidia ecosystem.
Trade pressSlicast · October 30, 2024 · Global · Source: techrepublic.com
importance 78

Google Cloud announced on October 30 at the App Day & Infrastructure Summit that it will enhance AI cloud infrastructure with new TPUs and NVIDIA GPUs. The sixth-generation Trillium NPU, first announced in May, is now in preview for cloud customers and powers many of Google Cloud's most popular services, including Search and Maps. "Through these advancements in AI infrastructure, Google Cloud empowers businesses and researchers to redefine the boundaries of AI innovation," said Mark Lohmeyer, VP and GM of Compute and AI Infrastructure at Google Cloud. "We are looking forward to the transformative new AI applications that will emerge from this powerful foundation."

The sixth-generation Trillium NPU delivers training, inference, and delivery of large language model applications at 91 exaflops in one TPU cluster. The new version offers a 4.7-times increase in peak compute performance per chip compared to the fifth generation, doubles the High Bandwidth Memory capacity, and doubles the Interchip Interconnect bandwidth. Trillium infrastructure meets the high compute demands of large-scale diffusion models like Stable Diffusion XL and can link tens of thousands of chips, creating what Google Cloud describes as "a building-scale supercomputer." Deniz Tuna, head of development at mobile app development company HubX, reported: "We used Trillium TPU for text-to-image creation with MaxDiffusion & FLUX.1 and the results are amazing! We were able to generate four images in 7 seconds — that's a 35% improvement in response latency and ~45% reduction in cost/image against our current system!"

In November, Google will add A3 Ultra VMs powered by NVIDIA H200 Tensor Core GPUs to their cloud services. The A3 Ultra VMs run AI and high-powered computing workloads on Google Cloud's data center-wide network at 3.2 Tbps of GPU-to-GPU traffic and will be available through Google Cloud or Google Kubernetes Engine. The Titanium ML network adapter uses NVIDIA ConnectX-7 hardware and Google Cloud's data-center-wide 4-way rail-aligned network to deliver 3.2 Tbps of GPU-to-GPU traffic, enhancing Google Cloud's optical circuit switching network fabric.

Hypercompute Cluster contains A3 Ultra VMs and can be configured via an API call, leverages reference libraries like JAX or PyTorch, and supports open AI models like Gemma2 and Llama3 for benchmarking. Google Cloud customers can access Hypercompute Cluster with A3 Ultra VMs and Titanium ML network adapters in November. Additionally, Google Cloud is preparing racks for NVIDIA's upcoming Blackwell GB200 NVL72 GPUs, which are anticipated for adoption by hyperscalers in early 2025 and will connect to Google's Axion-processor-based VM services.

Read the original
Google Cloud announces sixth-generation TPUs… · Slicast