Friday, August 7, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeData CentersReport
Data Centers · Report

Google Jupiter network fabric scales to 13 petabits per second, advancing datacenter interconnect capacity.

Hyperscaler network breakthrough enables higher GPU utilization and larger cluster sizes; drives infrastructure architecture evolution.
Trade pressSlicast · November 3, 2024 · Global · Source: cloud.google.com
importance 85

Google's network infrastructure has evolved significantly over 25 years, guided by a vision to handle exponential growth in user base and service demand. The current fifth-generation Jupiter data center network architecture now scales to 13 Petabits/sec of bisectional bandwidth, with hundreds of Jupiter fabrics deployed worldwide supporting hundreds of services, billions of active daily users, all Google Cloud customers, and some of the largest ML training and serving infrastructures in the world. To contextualize this capacity, the network could support a video call at 1.5 Mb/s for all 8 billion people on Earth.

This evolution has been guided by four key principles: predictable low latency through bandwidth headroom provisioning and maintaining 99.999% network availability; software-defined and systems-centric architecture leveraging SDN to qualify and globally release dozens of new features every two weeks; incremental evolution and dynamic topology enabling granular network refreshes and adaptation to changing workload demands through optical circuit switching and SDN; and traffic engineering with application-centric QoS to tailor the network to each application's needs. As the foundation of reliability for all other compute services, Jupiter networks deliver a factor of 50x more reliability than prior versions of Google's data center networks, with every bad minute rigorously defined and monitored across hundreds of clusters and millions of ports.

A seminal paper detailed how Jupiter data center networks scaled to 1.3 Pb/s of aggregate bandwidth by leveraging merchant switch silicon, Clos topologies, and Software Defined Networking. In 2022, Google announced that its Jupiter networks had scaled to over 6 Pb/s, with deep integration of optical circuit switching (OCS), wave division multiplexing (WDM), and the highly scalable Orion SDN controller. These technologies unlocked incremental network builds, enhanced performance, reduced costs, lower power consumption, dynamic traffic management, and seamless upgrades. Today, Jupiter supports native 400 Gb/s link speeds in the network core, with aggregation blocks consisting of 512 ports of 400 Gb/s connectivity for an aggregate of 204.8 Tb/s of bidirectional non-blocking bandwidth per block, and 64 such blocks delivering a total bisection bandwidth of 13.1 Pb/s—infrastructure powering Google's production data centers for over a year.

Looking forward, Google is charting the course for next-generation network infrastructure to support the age of AI. Teams are working on networking infrastructure for upcoming A3 Ultra VMs featuring NVIDIA ConnectX-7 networking, supporting non-blocking 3.2 Tbps per server of GPU-to-GPU traffic over RoCE (RDMA over converged ethernet), and future offerings based on NVIDIA GB200 NVL72. Over the coming years, Google plans to deliver significant advances in network scale and bandwidth per-port and network-wide, push the boundaries of end-host integration including transport and congestion control stacks, implement real-time topology engineering with deeper integration into compute and storage stacks, and refine host-based load balancing techniques to further enhance network reliability and latency.

Read the original
Google Jupiter network fabric scales to 13… · Slicast