Friday, September 11, 2026
AI 인프라 · 뉴스 & 분석
반도체·하드웨어리포트
반도체·하드웨어 · 리포트

NVIDIA의 Vera Rubin NVL72 GPU는 메가와트당 30배 높은 처리량과 GB300 NVL72보다 35배 낮은 토큰 비용을 제공합니다

NVIDIA 공식 — 개발 계획/제품 직접 확인
공식 공시Slicast · September 11, 2026 · 미국 · 출처: NVIDIA Blog

Agentic AI workloads require fundamentally different infrastructure than traditional chat or summarization tasks. While simple requests operate within 1K to 8K tokens, agentic sessions accumulate context across multiple steps and can reach hundreds of thousands of input tokens. An AI agent researching a company for investment decisions might query financial databases, search news and filings, invoke sub-agents for peer comparisons and valuations, then synthesize findings into a recommendation. Each step's accumulated tokens become input to the next, making long-context handling central to performance. This pattern repeats across use cases from software development to customer service to deep research.

Measured performance using the SemiAnalysis AgentX workload, which consists of recorded real-world agentic coding sessions with actual context growth, tool calls and sub-agent spawning preserved, shows Vera Rubin NVL72 delivering up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads. The Blackwell platform leads across multiple agentic models including Kimi K3, MiniMax M3, GLM5.3, Qwen 3.5 and DeepSeek V4 Pro. GB300 NVL72 itself delivers up to 15x better throughput per megawatt than the Hopper architecture on DeepSeek V4 Pro. For power-constrained AI factories, throughput per megawatt translates directly into more agentic work per energy footprint. Token cost drops to 35x lower than GB300 NVL72, enabling continuous agent operation at scale and determining profit margins on AI factory revenue.

Vera Rubin achieves these gains through extreme codesign spanning disaggregated serving that separates context processing from response generation, rate matching that synchronizes prefill and decode GPU speeds, large-scale expert parallelism distributing mixture-of-experts models, distributed KV-caching extending memory across the GPU domain with optional offloading to host and storage, KV-aware routing directing requests to GPUs holding relevant cached context, and fused CUDA kernels combining computation and communication into single execution passes. The enhanced fifth-generation Tensor Cores and third-generation Transformer Engine accelerate both prefill and decode stages. NVFP4 quantization compresses model weights to 4-bit precision. The NVL72 scale-up domain with sixth-generation NVLink and NVLink Switches delivers 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet. NVIDIA's software stack including TensorRT LLM and Dynamo is codesigned with hardware to enable these optimizations.

The full Vera Rubin platform is a seven-chip architecture including the Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX and ConnectX-9 SuperNIC, all purpose-built for AI factories deploying agents at scale. Vera Rubin is in full production and scaling across the ecosystem.

원문 보기