Friday, August 7, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Nvidia revealed Vera Rubin GPU specifications with 288GB memory, 22TB/s bandwidth, and 50 PFLOP/s AI throughput.

New capability metrics set competitive bar for infrastructure designs requiring massive memory and bandwidth for frontier AI models.
Trade pressSlicast · March 16, 2026 · Global · Source: wccftech.com
importance 80

NVIDIA has officially unveiled its next-generation AI data center platform called Vera Rubin, powered by the Rubin GPU and Vera CPU architectures. The platform is designed with a total of seven chips and six different racks, each serving a singular purpose to power next-generation AI datacenters. The Vera Rubin Compute Tray features a new mounting system that enables AI data centers to complete installation in just 2 hours instead of 2 days. The entire tray is liquid-cooled by hot water at 45 °C, alleviating pressure on data center infrastructure. This primary compute tray houses the new Rubin GPUs, featuring two massive reticle-sized dies and 8 HBM sites.

Each NVIDIA Rubin GPU features 288 GB of HBM4 memory, offering up to 22 TB/s of total bandwidth and 50 PFLOPs of NVFP4 compute performance. Each chip packs 336 billion transistors, with an additional 2.5 trillion transistors from the HBM4 memory. The Vera CPU offers extremely high single-threaded core performance, incredibly high data output, and extreme levels of energy efficiency, making it the world's first and only data center CPU to utilize LPDDR5 memory. NVIDIA will ship Vera CPUs standalone, with the company expecting this to open another multi-billion-dollar business front. Additional platform components include the NVLink Switch Tray with 6th Gen NVLINK, the Groq 3 LPX compute tray comprised of 8 GROK LPUs offering 500 MB of SRAM and 150 TB/s of SRAM bandwidth, and the Spectrum-X CPO Switch, the world's first co-packaged optics switch made at TSMC using NVIDIA's Cu-Litho technology.

The NVIDIA Vera Rubin NVL72 system offers a 10x performance per watt increase, 3.6 ExaFlops of NVFP4 performance, 1.6 PB/s of HBM4 bandwidth, and 260 TB/s of NVLINK6 interconnect speeds. NVIDIA Vera CPUs are also available in a 256 Vera CPU rack configuration, offering 300 TB/s of LPDDR5X bandwidth with all units connected using an ETL Spine, delivering 6.5x the throughput versus the previous-generation solution.

Vera Rubin-based products will be available from partners beginning in the second half of 2026. Leading cloud providers including Amazon Web Services, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure, along with NVIDIA Cloud Partners CoreWeave, Crusoe, Lambda, Nebius, Nscale, and Together AI, will offer these systems. Global system manufacturers Cisco, Dell Technologies, HPE, Lenovo, and Supermicro are expected to deliver a wide range of servers based on Vera Rubin products, alongside Aivres, ASUS, Foxconn, GIGABYTE, Inventec, Pegatron, Quanta Cloud Technology, Wistron, and Wiwynn.

AI labs and frontier model developers including Anthropic, Meta, Mistral AI, and OpenAI are looking to use the NVIDIA Vera Rubin platform to train larger, more capable models and to serve long-context, multimodal systems at lower latency and cost than with prior GPU generations.

Read the original
Nvidia revealed Vera Rubin GPU specifications… · Slicast