Friday, September 11, 2026
AI 인프라 · 뉴스 & 분석
반도체·하드웨어리포트
반도체·하드웨어 · 리포트

NVIDIA의 Vera Rubin NVL72 플랫폼이 코디자인 칩으로 10배 향상된 처리량을 제공하며 본격 생산으로 진입하고 있습니다

NVIDIA 공식 — 로드맵/제품의 직접 확인
공식 공시Slicast · August 11, 2026 · 미국 · 출처: NVIDIA Blog

NVIDIA Vera Rubin is ramping production with racks now running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius. The platform spans over 350 factory sites across 30 countries and represents the largest, most mature rack-scale supply chain assembled to meet AI compute demand.

Vera Rubin is built from chip to grid to deliver the highest performance per watt and the lowest token cost. CoreWeave's first benchmark on DeepSeek-R1 demonstrated the platform's capabilities: 10x more throughput per megawatt than Grace Blackwell NVL72, which CoreWeave says lands directly on the metric that matters most for power-constrained AI factories.

The platform's power comes from extreme codesign across seven integrated chips and five rack trays: Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX and Vera BlueField-4 STX, all engineered as a single system rather than assembled from separate off-the-shelf products. The NVIDIA Vera CPU sits at the platform's center, designed specifically for the agent era. Its custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth and 40 percent lower memory latency compared to competing chiplet designs, making it the most efficient single-threaded CPU for agentic workloads.

Networking performance has been dramatically improved. The platform's sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency and 10x higher packet rates than off-the-shelf Ethernet. For scale-out networking, Spectrum-X Ethernet combines 102.4 terabit Spectrum-6 switch systems and 1.6 terabit ConnectX-9 SuperNICs with adaptive routing, advanced congestion control, telemetry and open operating system support, enabling 1.6x higher RDMA bandwidth than off-the-shelf Ethernet. Leading infrastructure builders including CoreWeave, Microsoft, SpaceXAI and Tesla are deploying Spectrum-6 switches to accelerate their AI factories. NVIDIA's photonics with co-packaged optics for scale-out, the industry's first in volume manufacturing, adds 5x lower power and 10x higher mean time between interruptions versus pluggable transceivers, with CoreWeave, Lambda and Oracle Cloud Infrastructure among first adopters.

Hardware design innovations significantly improve manufacturing and operations. Three generations of rack-scale codesign have produced a Vera Rubin NVL72 system with no cables, fans or hoses in the tray, cutting compute tray assembly time from hours to one minute. The platform uses a 45-degree Celsius liquid cooling inlet design enabling chiller-free dry-cooler operation. For new AI factories, this higher-temperature dry cooling combined with the closed-loop liquid cooling system saves millions of gallons of water per megawatt annually.

Vera Rubin is becoming the foundation for Europe's next-generation AI infrastructure through an expanded Microsoft and Mistral partnership. A new multibillion-dollar agreement focuses on expanding AI infrastructure in Europe, with Mistral adding GPU capacity by deploying thousands of the latest NVIDIA Vera Rubin GPUs to increase AI compute availability and provide a shared platform for training, inference and large-scale deployment. This addresses Europe's need for AI infrastructure running under its own laws, close to home and fully within its control. Agentic systems can consume up to 15x more tokens than traditional AI applications, making efficient infrastructure a strategic priority that requires not just computing capacity but also data control, governance, resilience and strategic autonomy.

Compared with NVIDIA's GB200 NVL72, Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and costs one-tenth per million tokens, providing more intelligence within the same power footprint. Mistral Medium 3.5 and OCR 4 are now available in Microsoft Foundry, with Mistral models integrated into Microsoft Copilot Studio. Through Azure Local and Foundry Local, customers can use the same models, tools and operating patterns across cloud and customer-controlled environments.

CoreWeave was the first AI cloud to bring up and validate Vera Rubin NVL72 hardware and share measured performance numbers. Running a DeepSeek-R1 benchmark on Vera Rubin NVL72, CoreWeave observed 10x improvement in tokens per second per megawatt compared with Grace Blackwell NVL72. DeepSeek-R1's mixture-of-experts architecture makes all-to-all GPU communication critical at scale, and Vera Rubin NVL72's 260 terabyte per second all-to-all NVLink 6 fabric removes that constraint. CoreWeave deployed the NVIDIA Spectrum-X Ethernet SN6600-LD as switching fabric, built on the 102.4 terabit Spectrum-6 switch chip with liquid-cooled design, delivering 1.64 petabytes per second per rack with 100 percent more capacity than previous-generation air-cooled switches.

Google Cloud has deployed Vera Rubin NVL72 in its first A5X instance for London startup Ineffable Intelligence, which develops next-generation intelligent superlearner systems that continuously learn through operation.

원문 보기