NVIDIA has expanded its NVLink Fusion platform with NVHBM, a next-generation high-bandwidth memory technology that impro
The next wave of artificial intelligence is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, system performance relies on how compute, memory, storage, networking, and software are designed together. To help hyperscalers and AI innovators build next-generation semi-custom AI infrastructure, NVIDIA has expanded NVLink Fusion with NVHBM, a next-generation high-bandwidth memory technology that delivers higher memory performance and efficiency to XPUs. Leading memory partners will validate and offer this technology, extending advanced memory capabilities to NVLink Fusion customers.
Traditional HBM architectures place the memory controller on the XPU die, consuming valuable silicon area. NVHBM integrates NVIDIA’s custom memory controller directly into the HBM base die. This architectural shift delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption compared to standard HBM4E, while freeing up to 25% more area on the XPU compute die. By establishing a standard NVHBM implementation available from multiple memory providers, NVIDIA reduces the engineering effort required to integrate and qualify memory across suppliers, giving customers a faster path to market for custom AI chips.
Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion. The partnership will focus on enhancing performance and efficiency for AI workloads through NVHBM technology and the NVLink scale-up architecture. This builds on AWS’s previously announced support for NVLink Fusion, with Annapurna Labs supporting the platform on its next-generation Trainium chips starting with Trainium4. This integration will allow Amazon chips and NVIDIA GPUs to operate together within a common rack-scale architecture. Nafea Bshara, vice president of Annapurna Labs at Amazon, stated, “NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency. We look forward to this technology collaboration to benefit future AWS infrastructure designs.”
NVLink Fusion enables partners to connect custom XPUs and CPUs to NVIDIA’s rack-scale platform. Partners gain access to NVIDIA NVLink chiplets, NVLink-C2C, NVLink Switches, and NVIDIA MGX systems and racks, alongside a broad ecosystem of CPU partners, ASIC designers, system manufacturers, and technology providers. Offered with each generation of NVIDIA’s rack-scale system architecture, NVLink Fusion allows hyperscalers and AI-native companies to concentrate engineering resources on XPU innovation while leveraging a proven technology stack for scale-up and scale-out networking, rack-scale systems, and software. This creates a faster, lower-risk path to deploying semi-custom AI infrastructure.