NVIDIA reports that the Vera Rubin architecture delivers a 30-fold improvement in throughput per watt over Blackwell and cuts token processing costs by 35 times for agentic AI tasks.
NVIDIA’s Vera Rubin platform delivers disruptive token throughput at significantly lower costs than Blackwell, underscoring its capabilities in agentic AI workloads. These on-silicon performance metrics were validated by NVIDIA using real-world agentic coding trajectories. The evaluation relied on the SemiAnalysis AgentX benchmark, which assesses AI infrastructure—including Vera Rubin—across agentic-coding inference tasks using models such as Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro.
Initial benchmarks highlight the NVIDIA Blackwell architecture’s gains over previous generations. A GB300 NVL72 “Grace Blackwell” server achieves 15 times higher throughput per megawatt than H200 NVL8 “Hopper” solutions when running DeepSeek-v4-PRO 1.6T. Within identical power budgets, Blackwell sustains higher and more responsive agentic inference throughput. Additionally, Blackwell reduces token costs by 10 times per million tokens compared to its predecessor. This improved total cost of ownership enables AI operators to expand agent capacity within existing power and infrastructure limits, or maintain current capacity while further reducing operational expenses.
Blackwell’s performance scales exceptionally well with larger models. When tested with Kimi K3 2.8T, the Blackwell solution delivered an 80-fold increase in throughput per megawatt versus Hopper, while sustaining 215 tokens per second per user for interactivity—a level substantially beyond Hopper’s capabilities.
Building on these foundations, the NVIDIA Vera Rubin NVL72 platform advances agentic AI inference even further. It delivers a 30-fold increase in throughput compared to the Grace Blackwell NVL72 in DeepSeek-v4-PRO 1.6T workloads. In terms of user interactivity, Vera Rubin maintains approximately 160 tokens per second per user and peaks at roughly 280 tokens per second, comfortably surpassing Blackwell’s peak of under 180 tokens per second per user.
Vera Rubin also drastically reduces inference costs in agentic coding environments, offering a 35-fold decrease in cost per million tokens relative to Blackwell. This efficiency allows NVL72 deployments to run multiple agents continuously at scale across diverse workloads. Supported by NVIDIA DSX MaxLPS technology—which dynamically manages power distribution across GPUs, racks, and workload levels—AI factories can now provision up to 40 percent more GPUs within the same megawatt budget.
Current benchmarks focus exclusively on Vera Rubin silicon and do not yet account for Vera CPU performance in tool-calling operations, nor do they fully represent the complete “Extreme Codesign” seven-chip platform. As additional Vera Rubin systems come online, performance metrics for AI agents and inference are expected to improve further. Meanwhile, NVIDIA has officially commenced production across its entire AI portfolio, including Vera CPUs, Rubin GPUs, Vera Rubin servers, Groq 3 LPX chips, and its full suite of networking technologies.
Hassan Mujtaba, Senior Editor for hardware at Wccftech, brings years of industry experience specializing in deep-dive technical analysis of next-generation CPU and GPU architectures, motherboards, and cooling solutions. His reporting combines breaking news on emerging technologies with extensive hands-on reviews and benchmarking. For continued coverage, readers can follow Wccftech on Google News.