Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

NVIDIA states that the Vera Rubin NVL72 rack delivers up to 30 times more Agentic AI throughput per megawatt and reduces token cost by 35 percent compared to prior generations.

These efficiency gains directly improve economics for GPU cloud operators, enabling higher utilization rates and faster payback periods on multi-gigawatt campus deployments.
Trade pressSlicast · August 25, 2026 · US · Source: Google News
importance 90

NVIDIA states that its Vera Rubin NVL72 platform delivers up to 30 times greater agentic AI inference throughput per megawatt compared to the GB300 NVL72, while cutting the cost per million tokens by as much as 35 times.

These performance gains address a rapidly emerging infrastructure bottleneck: AI agents consume significantly more compute than conventional chatbot interactions. According to OpenRouter data cited by NVIDIA, agentic workloads require approximately 15 times more tokens than a standard chat query.

Unlike traditional models that generate responses from a single prompt, AI agents iteratively query databases, search documents, invoke tools, spawn sub-agents, analyze outputs, and sustain reasoning until a task concludes. Each step appends new context to the session, potentially expanding individual agent runs to hundreds of thousands of input tokens.

This operational pattern elevates the importance of long-context processing, memory management, and token-generation efficiency as enterprises deploy agentic applications at scale.

NVIDIA benchmarked the Vera Rubin NVL72 using the SemiAnalysis AgentX workload, which preserves real-world agentic coding sessions—including authentic context expansion, tool invocations, and sub-agent spawning. Running the DeepSeek V4 Pro model, the platform achieved up to 30 times higher throughput per megawatt than the GB300 NVL72.

These initial benchmarks remain pending independent review by SemiAnalysis and currently exclude Vera CPU performance for tool-calling operations.

NVIDIA’s prior-generation Blackwell platform already demonstrated substantial efficiency gains for comparable workloads. The GB300 NVL72 delivers up to 15 times greater throughput per megawatt than the Hopper architecture when running DeepSeek V4 Pro.

Vera Rubin extends these efficiency gains across the entire performance curve. For power-constrained data centers, NVIDIA notes that the enhanced throughput-per-megawatt ratio enables operators to run substantially more agentic workloads within existing power envelopes.

As power availability solidifies as a primary constraint on new AI infrastructure, these economic advantages grow critical. NVIDIA also confirmed that the Vera Rubin NVL72 reduces the cost per million tokens by up to 35 times relative to the GB300 NVL72.

Reduced token costs make continuous agent operation economically viable across extended-reasoning workflows such as software development, customer support, financial analysis, and research. NVIDIA achieved these gains through cross-layer co-design encompassing GPUs, networking, memory architecture, inference software, and workload orchestration.

Key architectural techniques include disaggregated serving, which decouples context processing (prefill) from response generation (decode) to enable independent scaling. Rate matching subsequently synchronizes prefill and decode GPU speeds to maximize utilization. For mixture-of-experts models, large-scale expert parallelism distributes individual expert networks across the NVL72 scale-up domain.

Distributed KV caching expands available memory across the GPU domain, while KV-cache offloading shifts inactive context to host memory or storage, eliminating the need to recompute previously processed data. Additionally, KV-aware routing directs incoming requests to GPUs already holding relevant cached context, minimizing redundant computation during prolonged agent sessions.

Fused CUDA kernels, including MegaMoE, merge computation and inter-GPU communication tasks to keep processors fully utilized rather than idle during data transfers.

The Rubin GPUs feature enhanced fifth-generation Tensor Cores and NVIDIA’s third-generation Transformer Engine to accelerate both prefill and decode phases. NVFP4 quantization compresses model weights to 4-bit precision, lowering memory demands and boosting throughput while preserving output fidelity, according to NVIDIA.

The NVL72 architecture features a large-scale GPU domain engineered to support the high-bandwidth, low-latency communication essential for expert parallelism and distributed KV caching. NVIDIA reports that its sixth-generation NVLink and NVLink Switch technologies deliver 10 times higher packet rates and three times lower latency than commercial Ethernet alternatives.

This hardware integrates with NVIDIA’s comprehensive inference software stack, featuring optimized CUDA kernels, TensorRT LLM, and the NVIDIA Dynamo serving framework.

Advanced power management further supports the efficiency strategy. NVIDIA’s DSX MaxLPS technology orchestrates power distribution across GPU, rack, and workload tiers, enabling operators to provision up to 40 percent additional GPUs within an identical megawatt budget.

Rather than representing a mere GPU iteration, Vera Rubin is architected as a seven-chip AI factory platform. The complete ecosystem comprises the NVIDIA Vera CPU, Groq 3 LPU, Rubin GPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX, and ConnectX-9 SuperNIC.

NVIDIA confirms that Vera Rubin has entered full production and is actively scaling across its hardware and infrastructure partner network.

As agentic AI transitions from experimental pilots to large-scale production, NVIDIA is shifting industry focus toward throughput per megawatt and token cost as primary performance metrics, moving beyond reliance on traditional single-request inference benchmarks.

Thank you for visiting Pulse 2.0. We work hard every day to bring leaders and decision makers like you the latest intelligence on business, finance, capital markets, deal flow, law, tech, and AI. Click here to subscribe to the Pulse 2.0 Newsletter.

Read the original
NVIDIA states that the Vera Rubin NVL72 rack… · Slicast