Monday, August 17, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeHeadlinesReport
Headlines · Report

Nvidia Rubin Ultra targets 768GB of HBM4E memory with the Kyber platform remaining on schedule.

Confirms next-generation GPU memory roadmap; 768GB capacity critical for long-context LLM inference and large model training demand.
Trade pressSlicast · August 14, 2026 · US · Source: Google News
importance 82

Nvidia's Rubin Ultra GPU is targeting up to 768GB of HBM4E memory using 12-Hi stacks, marking a substantial upgrade from the 288GB of HBM4 supported by current Rubin GPUs. The original design called for 1TB of memory using 16-Hi HBM4E stacks, but the target has been revised downward in response to anticipated supply constraints across the HBM4E production chain. Nvidia also is evaluating lower-capacity 8-Hi HBM4E configurations as a fallback option.

The Kyber rack-scale platform, which includes the NVL144 and NVL576 systems, remains active and on schedule for a second-half 2027 release alongside Rubin Ultra. Reports had circulated suggesting that Rubin Ultra could slip into 2028, which would represent a meaningful setback in Nvidia's annual GPU release cadence. Nvidia has publicly denied those claims and reaffirmed the second-half 2027 target.

The company's roadmap positions Vera Rubin GPUs for 2026, followed by Rubin Ultra in 2027. Rubin Ultra incorporates a multi-chiplet design intended to increase GPU density per rack, with the Kyber NVL576 scaling to 576 GPUs in a single rack configuration.

Nvidia sources HBM from SK Hynix, Samsung, and Micron. SK Hynix is widely regarded as the leading supplier of HBM at current node generations, and its production capacity for HBM4E will be a gating factor for the entire industry in 2026 and 2027. Samsung and Micron are ramping their own HBM4E programs but have faced qualification delays with major customers. Nvidia's multi-supplier approach reduces single-source dependency, though it does not eliminate category risk entirely.

The memory upgrade carries significant practical weight. In large language model training, a GPU's memory bandwidth and capacity determine how large a model can fit on a single chip and how efficiently the system can train without constantly offloading data between chips. A jump from 288GB to 768GB reduces inter-chip data transfers, enabling faster training at lower energy cost—far more than a specification-sheet metric.

Read the original
Nvidia Rubin Ultra targets 768GB of HBM4E… · Slicast