Wednesday, September 16, 2026
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

NVIDIA reports token-per-watt optimization advances for Vera Rubin and DSX platforms at AI Infra Summit, highlighting power efficiency as competitive metric.

Efficiency (tokens per watt) as published metric signals datacenter density is power-constrained, not compute-constrained; affects facility planning.
NewswireSlicast · September 15, 2026 at 17:01 UTC · US · Source: NVIDIA Blog
importance 75

Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, addressed AI factory efficiency at the AI Infra Summit in Santa Clara, where attendance swelled to over 8,000 this year from 3,500 last year. Buck highlighted new platform collaborations and partnerships advancing energy efficiency in AI infrastructure.

Amazon's Annapurna Labs is collaborating with NVIDIA on NVHBM custom high-bandwidth memory technology. d-Matrix is integrating with NVLink Fusion to pair NVIDIA Vera CPUs with d-Matrix Raptor XPUs for ultra-low-latency inference at scale. Emerald AI and NVIDIA demonstrated a commercial AI factory flexible-load program with Silicon Valley Power. Lambda improved performance per watt by 23% using NVIDIA DSX MaxLPS. Pinterest is deploying the NVIDIA Blackwell platform with NVIDIA Dynamo inference software to add conversational AI to visual discovery.

NVIDIA's response to the rising demand from agentic AI workloads is a full-stack platform spanning Vera Rubin systems, Dynamo inference software, NeMo libraries, and NVIDIA networking—including NVLink for scale-up computing, Spectrum-X Ethernet, ConnectX SuperNICs for thousands-node deployments, BlueField-powered context-memory storage, and BlueField DPUs for infrastructure security.

The industry's performance metric is shifting from peak performance to validated agentic tokens per megawatt, requiring AI factories to be codesigned from silicon to grid. NVIDIA DSX MaxLPS delivers up to 1.4x more tokens per megawatt through factory-wide power optimization, while NVLink unites large-scale accelerated computing into a single high-performance system—enabling customers to generate more tokens, improve efficiency, and extract maximum value from every megawatt deployed.

Silicon Valley Power's flexible-load interconnection program enables AI factories to support grid flexibility. Emerald AI, working with NVIDIA, demonstrated automated load reduction while protecting workload performance, responding to hundreds of demand signals. Using NVIDIA DSX Flex, Emerald AI's Conductor grid-responsive power management software dynamically adjusts energy consumption based on real-time grid signals and hybrid energy sources. DSX Flex receives load-shedding requests, demand-response events, and pricing signals, then automatically acts within a predefined workload hierarchy—protecting critical jobs while pausing lower-priority tasks temporarily, then resuming. This allows facilities to participate in grid efficiency while reducing demand when needed and unlocking additional capacity for growth.

Lambda released validation results for NVIDIA DSX MaxLPS on Blackwell servers. The software continuously monitors power consumption across GPUs and racks, dynamically shifting available power and reclaiming capacity left unused by static provisioning. Running 19 nodes within the power budget typically allocated to 16 full-power nodes, Lambda increased cluster-wide token throughput by 24%—from approximately 4 million to 5 million tokens per second—while improving performance per watt by 23%. For next-generation NVIDIA Vera Rubin NVL72 deployments, DSX MaxLPS can enable up to 40% more GPU capacity within the same megawatt budget in suitable environments, allowing operators to expand AI capacity and token throughput within existing power envelopes.

Power represents the primary constraint in AI factories, making tokens per megawatt the critical efficiency metric. NVIDIA's full-stack AI factory built on Vera Rubin NVL72 integrates systems, networking, software, and power management. At the factory level, DSX MaxLPS dynamically shifts power across racks as demand fluctuates. Inside each rack, Intelligent Power Smoothing software and expanded energy buffering absorb short spikes, letting systems run closer to sustained demand and converting stranded headroom into productive compute. Results include up to 40% more GPUs within the same site-power envelope and up to 35% higher token throughput without new power lines.

For agentic AI—where agents chain reasoning steps and tool calls—latency and context length multiply quickly. NVIDIA Groq 3 LPX adds deterministic ultralow-latency inference to Vera Rubin, complementing DSX MaxLPS. The combined platform delivers up to 35x higher token throughput per megawatt than GB200 NVL72 for 2-trillion-plus-parameter models at long context. On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX achieved 2,529 output tokens per second per user—headroom enabling agents to run more reasoning steps and tool calls within the same response budget even as workloads scale.

SemiAnalysis AgentX measures inference on recorded real-world agentic coding sessions, preserving actual context growth, tool call delays, and sub-agent spawning rather than relying on single-request benchmarks. NVIDIA Vera Rubin NVL72 results are now live on the AgentX dashboard. On the DeepSeek V4 Pro model, Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72. Agentic workloads differ fundamentally from traditional request-response inference: a single session accumulates hundreds of thousands of input tokens as agents reason, call tools, and spawn sub-agents—roughly 15x the token volume of a simple chat—with wide variability in input and output length. AgentX captures this end-to-end behavior, complementing standardized suites like MLPerf Inference. The 30x throughput-per-megawatt gain means 30x more agentic work from the same energy footprint; AgentX results show up to 45x lower cost per million tokens. In power-constrained deployments, throughput per megawatt determines revenue generation, and cost per million tokens determines margin. This performance stems from extreme codesign: the NVL72 scale-up domain, sixth-generation NVLink, NVFP4 precision on fifth-generation Tensor Cores, and an inference stack spanning TensorRT LLM and Dynamo. Vera Rubin is in full production and scaling ecosystem-wide, with performance continuing to improve through software optimization.

Startups are validating NVIDIA Vera CPU performance for agent workloads and demanding applications. Perplexity benchmarked the Vera CPU for its new SPACE secure sandbox platform for agentic AI, reporting 1.9x faster sandbox starts. Daytona tested the Vera CPU on agentic workloads, citing "serious gains for agentic execution."

Read the original
NVIDIA reports token-per-watt optimization… · Slicast