Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

NVIDIA detailed major upgrades to its BlueField-4 DPU and scale-in networking architecture at Hot Chips 2026.

This advances the disaggregated data center fabric, directly impacting how hyperscalers and neoclouds will interconnect GPU clusters for next-gen AI training.
Trade pressSlicast · August 26, 2026 · Global · Source: ServeTheHome
importance 85

At Hot Chips 2026, NVIDIA presented the BlueField-4 Processor, the fourth generation of its data center processing unit (DPU). The company positions the DPU as a critical component within its broader argument that agentic AI represents the most complex computing workload to date. By integrating the CPU, GPU, and network infrastructure, the BlueField-4 enables agents to execute the observe, reason, and act cycles efficiently.

The BlueField-4 operates within NVIDIA’s Vera Rubin platform, which the company describes as a full-stack AI factory spanning seven chips across five racks. Alongside the Vera CPU and Spectrum-6 switches, the DPU serves as a co-designed element embedded directly into every Vera Rubin system. This marks a strategic shift from treating the DPU as an optional PCIe expansion card to making it a foundational layer of the architecture. As hyperscalers deploy machines at massive scale, embedding DPUs directly into the NVIDIA AI Factory design aligns with industry demands for integrated data center infrastructure.

NVIDIA structures its networking strategy across distinct layers. Scale-up, managed through NVLink, allows multiple GPUs to function as a single system within a tenant, ensuring continuous GPU utilization within a node before traffic reaches the wider fabric. Scale-out leverages ConnectX and Spectrum-X switches to unify the data center into a single compute unit. Scale-in handles traffic between the host and the broader data center, often traversing multiple tenants and applications. Finally, storage-scale addresses high-performance data access for AI workloads.

To address scale-in performance, NVIDIA introduced the BlueField-4 Spectrum-X scale-in network, an end-to-end solution optimized for mixed agentic workloads. Initial benchmarks highlight storage access improvements: Spectrum-X Ethernet delivers 1.5x faster access for 5 GB files, 1.4x for 10 GB files, and 1.3x for 50 GB files compared to standard off-the-shelf Ethernet.

Modern DPUs require substantial network acceleration paired with robust compute capabilities. The BlueField-4 combines ConnectX-9-class networking with Grace CPU-class cores to handle control plane tasks, encryption, and high-throughput storage I/O. This contrasts sharply with earlier cloud DPUs designed for general-purpose computing, which treated the DPU as a replaceable add-on. Agentic AI demands a fundamentally different infrastructure tier, particularly regarding bandwidth. While legacy cloud DPUs operated in the 200 Gb/s range, the AI-focused BlueField-4 delivers 7 Tb/s of aggregate bandwidth. Furthermore, whereas older cards relied on PCIe slots, the BlueField-4 is engineered at the system level.

Hardware specifications for the BlueField-4 include a Grace CPU featuring 64 Arm Neoverse V2 cores clocked at 1.7 GHz, LPDDR5 bandwidth of 275 GB/s, and ConnectX-9 networking supporting 800G Ethernet across 200G PAM4 SerDes. The architecture incorporates inline encryption and a PCIe Gen6 x16 host link. The 1.7 GHz core frequency appears lower than configurations observed in GB300 AI systems and prior Grace servers, likely reflecting a deliberate optimization for a reduced power envelope, particularly given that only 64 Arm cores are active in this configuration. Performance comparisons indicate these V2 cores will significantly outperform earlier architectures such as the Xsight Labs E1 DPU, which utilized N2 cores.

The Vera Rubin platform presents extreme bandwidth requirements. Each compute tray demands 7 Tb/s of aggregate network bandwidth, constructed from four 1.6 Tb/s GPU scale-out links plus 800 Gb/s of scale-in connectivity routed to the DPU. Without a DPU owning the data path, NVIDIA argues that end-to-end GPU network security and tenant isolation become unachievable. While traditional approaches may isolate scale-in traffic, scale-out traffic remains exposed. Additionally, non-DPU architectures multiply power consumption, as each NIC requires a dedicated CPU, memory subsystem, and management connection—resulting in approximately 4x higher power draw compared to deploying a BlueField-4 ahead of the host.

The BlueField-4 addresses these challenges through Astra, delivering 7 Tb/s of CSP-grade security and management capabilities. By owning the data path, the DPU enforces zero-trust policies and collects telemetry at line rate without impacting host resources. Management functions utilize dedicated links, eliminating the need to place a BlueField-4 on every ConnectX-9 interface. Paper metrics confirm that measured bandwidth tracks ideal linear scaling as NIC counts increase, with encrypted and virtualized scale-out traffic remaining secure and monitored across multi-terabit aggregates. While competitors like Broadcom have demonstrated detailed performance views, the BlueField-4 occupies a distinctly different performance class compared to devices such as the Thor Ultra.

Live demonstrations showcased the 7 Tb/s platform-level DPU operating within a Vera Rubin system, illustrating how the BlueField-4 monitors all ConnectX-9 NICs and associated traffic. For GPU storage access, NVMe over Fabrics running on the BlueField-4 Grace variant delivers 1.6 Tb/s using eight cores and achieves 20 million IOPS with sixteen cores, providing a two-fold acceleration in data delivery to the GPUs. On the CPU side, the DPU enables transparent agentic security through in-silicon policy enforcement. DOCA Flow, DOCA Vault, and DOCA Argus map file, object, processor, and network contexts to maintain policy enforcement at agent speeds. The architecture also maintains line-rate forwarding and security checks for small-packet traffic, preventing control plane collapse under tiny frames.

Storage-Scale capabilities extend agent context across a tiered memory map. Utilizing DOCA MEMOS, the system achieves 3.2 Tb/s of storage access, delivering 10x IOPS and 5x efficiency by moving key-value context from GPU HBM through system memory to local and network storage. NVIDIA currently applies the “BlueField-4” designation to both the Grace-integrated and single ConnectX-9 variants, extending across single and dual Vera configurations with varying ConnectX-9 implementations. Industry feedback suggests that adopting distinct nomenclature, such as BlueField-4G and BlueField-4V, would improve clarity across these configurations.

NVIDIA positions the BlueField-4 as the first AI-native DPU, spanning scale-up, scale-out, scale-in, and storage-scale domains. The architecture marks a meaningful evolution: the DPU transitions from a peripheral add-in to the trusted front end of the AI factory. By leveraging a substantially more powerful CPU than previous generations, the BlueField-4 provides the necessary headroom to manage security and storage workloads that earlier iterations struggled to support. The result represents a substantial generational leap in data center infrastructure design.

Read the original
NVIDIA detailed major upgrades to its… · Slicast