NVIDIA announced at the AI Infra Summit a comprehensive shift in how AI infrastructure is evaluated — from peak performa
NVIDIA Vice President Ian Buck addressed over 8,000 attendees at the AI Infra Summit in Santa Clara on September 15, discussing how AI infrastructure is being redesigned around power efficiency and agentic workloads. The company introduced new collaborations and product announcements that collectively demonstrate a fundamental shift in the metrics used to evaluate data center performance.
Amazon's Annapurna Labs is collaborating with NVIDIA on NVHBM custom high-bandwidth memory technology. D-Matrix is integrating NVIDIA's NVLink Fusion platform with its Raptor XPUs. In the commercial space, Emerald AI and NVIDIA demonstrated an AI factory flexible-load program working with Silicon Valley Power, Lambda improved performance per watt by 23 percent with NVIDIA DSX MaxLPS, and Pinterest is using the NVIDIA Blackwell platform with Dynamo inference software for conversational visual discovery.
The underlying driver is agentic AI, which demands new infrastructure design. NVIDIA's full-stack platform spans Vera Rubin systems, Dynamo inference software, NeMo libraries and comprehensive networking solutions including NVLink for scale-up computing, Spectrum-X Ethernet and ConnectX SuperNICs. The company is promoting a shift from measuring peak performance to measuring validated agentic tokens per megawatt. DSX MaxLPS delivers up to 1.4 times more tokens per megawatt through factory-wide optimization, while NVLink enables large-scale accelerated computing to operate as a single system.
Silicon Valley Power operates a flexible-load interconnection program, and Emerald AI demonstrated automated load reduction while protecting workload performance. The system responded to hundreds of demand signals from the utility. Emerald AI plans to deploy DSX Flex with its Conductor grid-responsive power management software, which dynamically adjusts energy consumption based on real-time grid signals. DSX Flex receives load-shedding requests, demand-response events and pricing signals, then automatically pauses lower-priority jobs while maintaining critical workloads. This enables AI factories to function as flexible grid resources.
Lambda released results on NVIDIA Blackwell servers with DSX MaxLPS, running 19 nodes within the power budget normally allocated to 16 full-power nodes. Lambda achieved a 24 percent increase in cluster-wide token throughput, from approximately 4 million to 5 million tokens per second, while improving performance per watt by 23 percent. For next-generation Vera Rubin NVL72 factories, DSX MaxLPS can enable up to 40 percent more GPU capacity within the same megawatt budget.
In power-constrained environments, NVIDIA demonstrated how Vera Rubin NVL72 with Groq 3 LPX delivers up to 35 times higher token throughput per megawatt than GB200 NVL72 for models exceeding 2 trillion parameters at long context lengths. On a 100,000-token-context Qwen 3.8 27B workload, Groq 3 LPX achieved 2,529 output tokens per second per user. The Vera Rubin platform includes Intelligent Power Smoothing software and expanded energy buffering to absorb power spikes and run closer to sustained demand.
NVIDIA released performance results on the SemiAnalysis AgentX benchmark, which measures inference on recorded real-world agentic coding sessions with actual context growth, tool call delays and sub-agent spawning. Vera Rubin NVL72 delivers up to 30 times higher throughput per megawatt than NVIDIA GB300 NVL72 on the DeepSeek V4 Pro model. The company emphasized that agentic workloads differ fundamentally from single-request tasks, with sessions accumulating hundreds of thousands of input tokens as agents reason, call tools and spawn sub-agents. The efficiency gains translate to business impact: up to 30 times more agentic work from the same energy footprint and up to 45 times lower cost per million tokens. In power-constrained deployments, throughput per megawatt determines how much revenue an AI factory can generate.