Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Nvidia is raising AI server prices by approximately 15% due to soaring memory costs, which will add roughly $5 billion to the capital expenditure of a 1GW data center.

The DRAM and HBM supply squeeze directly inflates hyperscaler capex budgets and forces faster migration to memory-efficient architectures or alternative suppliers.
Trade pressSlicast · August 24, 2026 · China · Source: 量子位
importance 88

According to Bloomberg, several of NVIDIA’s largest customers have been notified that AI servers scheduled for delivery in early next year could face price increases exceeding 15%. The affected hardware spans the current Grace Blackwell architecture and the next-generation flagship Vera Rubin platform.

Citing two individuals who received notifications from server manufacturers, The Information reports that select GB300 and Vera Rubin 200 systems are projected to see price hikes of approximately 17%. Using the pricing model referenced in the report, constructing a 1-gigawatt (GW) AI data center could incur an additional cost of at least $5 billion from this single round of adjustments alone. At current estimates, a 72-GPU Vera Rubin rack costs roughly $7 million; following the increase, each unit could reach approximately $8 million.

Interviews with server vendors and GPU cloud providers previously revealed that certain components comprising NVIDIA AI servers can experience price fluctuations of up to 40% within a single week. Another executive noted that NVMe storage within GB300 servers was historically one of the most volatile segments; currently, some racks are priced 10% to 15% above their established baseline. Retail market data further shows median prices for the RTX 5060 Ti 16GB have surged by up to 39%, the RTX 5070 by 36%, the RTX 5060 by 27%, and the RTX 5090 by approximately 9%. Note that these figures reflect retail market pricing, influenced by supply-demand dynamics and channel inventory, and should not be conflated entirely with official NVIDIA price adjustments.

Professional GPUs have also experienced significant price movements. The 96GB VRAM RTX PRO 6000 Blackwell rose from approximately $8,565 in early 2025 to $13,250 in June 2026, and further to $16,000 by August 2026. Price pressure is cascading upward across NVIDIA’s entire product lineup, from consumer graphics cards and professional accelerators to multi-million-dollar data center racks.

TrendForce reports that starting in the third quarter of this year, NVIDIA expanded its evaluation scope for the Rubin Ultra beyond its initial focus on 12-Hi HBM4E to include multiple configurations such as 8-Hi HBM4E, 12-Hi HBM4, and 8-Hi HBM4. Final specifications remain undetermined due to persistent DRAM supply constraints heading into 2027, alongside uncertainties surrounding the validation timeline and mass-production yield rates for 12-Hi HBM4E. In late March, TrendForce forecasted that traditional DRAM contract prices would rise 58% to 63% quarter-over-quarter in Q2 2026. Server DRAM contract prices were similarly projected to increase by 13% to 18%. The firm further anticipates that server DRAM prices may continue to climb quarter-over-quarter from the second half of 2026 through the second half of 2027.

As AI server deployments scale, demand for High Bandwidth Memory (HBM) intensifies, prompting memory manufacturers to prioritize advanced process nodes and wafer capacity for HBM production. HBM requires larger dies, greater stacking layers, and more complex packaging processes, meaning identical wafer investments do not yield proportional bit outputs. TrendForce projects that server RDIMM bit supply will grow only 15% to 20% year-over-year in 2027, significantly lagging behind server CPU shipment growth. Since the first half of this year, several cloud service providers and server OEMs have begun substituting 96GB and 128GB RDIMMs with 32GB and 64GB modules in certain systems to mitigate procurement costs.

Due to LPDDR5X supply shortages, NVIDIA has decided to halve the SOCAMM memory capacity in its next-generation Vera Rubin Superchip. Based on preliminary capacity allocations from Samsung, SK Hynix, and Micron, suppliers can currently fulfill only about 60% of NVIDIA’s estimated LPDDR requirements. Over recent years, AI workloads have aggressively consumed GPUs, driving a surge in HBM demand, which in turn continues to squeeze conventional DRAM capacity. By 2026, memory prices are now feeding back to inflate GPU and AI server costs. However, rising memory expenses are merely the visible factor; underlying drivers include increased integration complexity in high-end system designs, broader supply chain constraints, and intense competition among cloud providers and large language model developers for faster delivery timelines. Critical infrastructure factors—particularly power availability, land acquisition, grid interconnection, thermal management, networking, and overall campus completion schedules—are now directly dictating when deployed GPUs can begin generating revenue.

On August 21, NVIDIA announced a minority equity investment in U.S. data center infrastructure developer Cloverleaf Infrastructure. Cloverleaf assists data center developers in securing suitable land, negotiating with utility companies and energy suppliers to secure power and other infrastructure resources, and converting these conditions into viable data center projects. Under the partnership, Cloverleaf will deploy NVIDIA’s DSX platform to optimize site selection, power, cooling, and compute infrastructure planning. Previously, Cloverleaf has partnered with institutions including Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR, aiming to mobilize over $500 billion in third-party capital for long-term AI infrastructure development. NVIDIA’s strategic concerns have expanded far beyond chip manufacturing to encompass customer purchasing power, data center siting, power sourcing, and post-delivery commissioning timelines.

Reports indicate that Amazon, Microsoft, Google, and Meta are all increasing investments in proprietary AI chips to reduce reliance on NVIDIA. In its latest published AVO research, an autonomous agent first optimizes GPU kernels before migrating them to the ARC-AGI-3 benchmark. Official disclosures reveal that AVO explored over 500 optimization directions, submitted 40 kernel versions, and achieved up to a 10.5% performance improvement over FlashAttention-4 on the DGX B200.

Read the original
Nvidia is raising AI server prices by… · Slicast