Nvidia announces Grace-Hopper hybrid CPU-GPU systems combining high memory bandwidth for AI workloads.
Nvidia is launching its Grace Arm-based CPU and associated superchip offerings in large-scale systems only, with no smaller 1U or 2U rack server configurations comparable to the DGX-A100 or DGX-H100 systems currently available. The Grace-Grace superchips combine two CPUs with NVLink 4.0 chip-to-chip links, while the Grace-Hopper GH200 superchips pair a Grace CPU with a Hopper GH100 GPU accelerator. The DGX GH200 system combines 256 Grace-Hopper superchips into a single shared memory GPU cluster, available in 32, 64, 128, and 256 GPU configurations. Reference designs called MGX have been proposed with these superchips, though Nvidia expects the bulk of initial shipments to go to the larger DGX systems given that demand will far exceed supply.
Nvidia's Jensen Huang announced the general availability of the Grace CPU during a keynote address at Computex in Taipei, Taiwan, which took place on a Monday morning in Asia. The company has not yet revealed an official product name for the Grace CPU or the Grace-Grace superchip, though based on Nvidia's naming conventions—which use "G" for GPU, the product codename initial, and "X00" for tier designation—the names would logically be "CG100" for a single Grace CPU and "CG200" for a Grace-Grace superchip. As of the article's publication, Nvidia had not released complete specifications or feeds and speeds for the Grace device, though whitepapers were being prepared for both the Grace chip and the new DGX GH200 system.
The DGX GH200 is not solely a product offering but a reference architecture that Google, Meta Platforms, and Microsoft plan to build upon as they develop AI training infrastructure to handle "1 trillion parameter large language models and massive recommendation systems." Nvidia's decision to create its own CPU stems from a longer history of strategic expansion in the datacenter. The company's ambitious "Project Denver" plan, announced in January 2011, proposed creating Arm-based CPUs to serve as complete replacements for X86 processors, combining powerful CPU and GPU capacity. That effort eventually ceased without public explanation. Decades later, with other companies having repeatedly failed to create server-class Arm CPUs, Nvidia restarted CPU development, particularly because IBM's Power10 chips could not effectively support the NVLink 4.0 protocol that Nvidia needed. With deep learning recommendation models and large language models increasingly requiring far more memory than GPU high-bandwidth memory alone could provide, Nvidia determined it needed direct control over its hardware stack and the ability to connect high-capacity memory faster than PCI-Express allowed.
The Grace CPU's defining feature is its use of LPDDR5 memory—the type found in laptops—which allows Nvidia to deliver 546 GB/sec of bandwidth across 32 channels to up to 512 GB of memory. This configuration provides 8X the capacity of HBM2e memory at approximately one third the bandwidth and one third the cost per GB, matching the price of DDR5 main memory while offering 4X the channels, one eighth the capacity, and 1.5X the bandwidth compared to DDR5. This memory capacity and bandwidth combination makes the Grace CPU suitable for both traditional HPC simulation and modeling as well as AI training for large language models and recommendation systems.