Lambda’s GPU cloud cluster has reached 5 million tokens per second throughput, leveraging NVIDIA’s DSX MaxLPS technology to achieve a 40% improvement in token efficiency per megawatt.
AI demand continues to place a significant strain on global energy infrastructure. To address this challenge, NVIDIA is deploying a comprehensive suite of energy efficiency optimizations across its full-stack AI factory platform. This ecosystem spans Vera Rubin systems, Dynamo inference software, NeMo libraries, and advanced networking solutions including NVLink for scale-up computing, Spectrum-X for Ethernet, ConnectX SuperNICs for connecting thousands of nodes, BlueField-powered context-memory storage, and BlueField DPUs for infrastructure security.
At the AI Infra Summit, NVIDIA highlighted how its DSX MaxLPS technology drives up to 40% higher token throughput per megawatt on Vera Rubin platforms. The software is already operational through Emerald AI’s Conductor platform, which demonstrates DSX’s grid orchestration capabilities. With DSX Flex, Conductor dynamically adjusts a data center’s power consumption based on real-time grid conditions. The system aims to reduce electricity demand during periods of grid constraint while ensuring zero interruptions for critical AI workloads. To date, the factory has processed 200 grid-condition signals flawlessly without any user intervention.
Cloud provider Lambda, which serves over 10,000 customers ranging from AI-native startups to hyperscalers, recently validated the platform’s performance. Testing was conducted on a five-rack, 19-node cluster. By operating all 19 nodes within the same power budget previously allocated to 16 nodes running at full power, Lambda achieved a 24% increase in cluster-wide token throughput—rising from approximately 4 million to 5 million tokens per second. Performance per watt improved by 23%, confirming the platform’s ability to deliver higher throughput at a fixed power budget.
These efficiencies directly translate to enhanced performance across Vera Rubin and Groq 3 LPX rack configurations. Facilities can now accommodate up to 40% more GPUs within the same site-power envelope and achieve up to 35% higher token throughput without requiring new power lines. In practical testing, a 100K-context Qwen 3.8 27B workload on Groq 3 LPX delivered 2,529 output tokens per second per user. This additional headroom enables AI agents to execute more reasoning steps and tool calls within the same response budget, even as workloads scale.
Beyond GPU optimizations, early feedback from startups worldwide highlights substantial performance gains with NVIDIA’s Vera CPUs. Independent benchmarks report up to 1.9x overall performance improvements, significantly reduced latencies, 73% higher throughput, and a 3.3x throughput increase in query tasks compared to competing processors. Perplexity evaluated the Vera CPU for its new SPACE secure sandbox platform designed for agentic AI, recording 1.9x faster sandbox initialization times. Daytona reported “serious gains for agentic execution,” while ClickHouse noted on its ClickBench analytical database benchmark that the Vera CPU was the fastest machine the team had ever measured, calling it a “strong signal of what’s ahead for CPU performance for data-intensive workloads.”
Further validation came from DeepInfra, which found the Vera CPU outperformed competitors across all tested metrics, including a 2.2x reduction in orchestration step latency. Prime Intellect observed that the architecture maintains high bandwidth and consistently low memory latency under parallel workloads—a critical requirement for predictable agentic AI performance. Redpanda recorded 5.5x lower latencies and 73% higher throughput versus other CPUs, Starburst demonstrated 3x faster query throughput, and Kinetica achieved 2.7x faster analytical query performance compared to traditional processors.
NVIDIA and its ecosystem partners are fundamentally redefining data centers as co-designed AI factories, shifting the industry standard from peak FLOPS to validated tokens per megawatt. This transformation is being driven by innovations ranging from Annapurna’s NVHBM work and d-Matrix’s NVLink Fusion integrations to Emerald AI’s grid-flexible load program with Silicon Valley Power, Lambda’s 23% performance-per-watt gain with DSX MaxLPS, Pinterest’s Blackwell-plus-Dynamo visual AI pipeline, Vera Rubin NVL72’s 30x AgentX throughput, Groq 3 LPX’s extreme latency advantages, the expanding wave of Vera CPU validations, and NVLink 6’s factory-scale resiliency. The overarching message remains consistent: integrated systems spanning silicon, software, and grid infrastructure can extract greater token yields, safeguard priority workloads, and maximize every megawatt. By converting constrained power into expanded capacity, these advancements are driving down cost per token and delivering sustained economic value for the next generation of AI infrastructure.