NVIDIA reports that its new Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt and up to 35x lowe
According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request because agents and sub-agents continuously reason, query databases, invoke tools, and synthesize information until a task is complete. Unlike chat or document summarization where sequences typically range from 1K to 8K tokens, agentic sessions accumulate context across steps and can reach hundreds of thousands of input tokens with wide variability in both input and output lengths. This makes long-context handling central to agentic AI performance across software development, customer service, and deep research, requiring infrastructure to evolve beyond measuring single inference requests.
New measured performance data demonstrates that NVIDIA Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 on agentic workloads. These early results were measured using the SemiAnalysis AgentX workload, which consists of recorded real-world agentic coding sessions with actual context growth, tool calls, and sub-agent spawning preserved. For power-constrained AI factories, this translates directly into 30x more agentic work for the same energy footprint. The Vera Rubin platform also extends the advantage seen in earlier architectures, where GB300 NVL72 delivers up to 15x better throughput per megawatt than the NVIDIA Hopper architecture on the DeepSeek V4 Pro model. On that same model, Vera Rubin lifts performance across the entire Pareto curve to achieve up to 30x higher throughput per megawatt than GB300 NVL72. These figures currently do not reflect Vera CPU performance for tool calling and remain pending SemiAnalysis review.
Throughput per megawatt directly determines AI factory revenue, while cost per million tokens dictates profit margins. At up to 35x lower cost per million tokens than GB300 NVL72, Vera Rubin NVL72 enables continuous, large-scale agent deployment. To maximize efficiency, NVIDIA DSX MaxLPS technologies manage power across the GPU, rack, and workload levels, provisioning up to 40% more GPUs within the same megawatt budget. The underlying hardware leverages enhanced fifth-generation Tensor Cores and third-generation Transformer Engines to accelerate both prefill and decode stages, alongside NVFP4 quantization that compresses model weights to 4-bit precision without sacrificing output quality. The NVL72 scale-up domain facilitates high-bandwidth, low-latency inter-GPU communication essential for large-scale expert parallelism and distributed KV-caching, powered by sixth-generation NVIDIA NVLink interconnect technology and NVLink Switches that deliver 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet alternatives.
This extreme codesign extends throughout the software stack, spanning optimized CUDA kernels, inference runtimes like NVIDIA TensorRT LLM, and serving frameworks like NVIDIA Dynamo. The complete seven-chip Vera Rubin platform also integrates the NVIDIA Vera CPU, Groq 3 LPU, NVLink 6 Switch, BlueField-4 DPU, Spectrum-6 SPX, and ConnectX-9 SuperNIC, all purpose-built for AI factories deploying agents at scale. While these results reflect current Vera Rubin NVL72 performance, continuous software optimizations will drive further improvements across both Vera Rubin NVL72 and GB300 NVL72. Vera Rubin is now in full production and scaling across the partner ecosystem, establishing a high-performance foundation for the next generation of agentic AI.