NVIDIA demonstrated that AI factories can dynamically adjust workloads in response to grid demand signals, recovering st
On an August evening in Silicon Valley, as air conditioning loads spiked during peak heat, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram watched on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers at the data center, and people from the utility itself. Nobody touched anything. Emerald AI's Conductor platform, a grid-orchestration system from NVIDIA partner Emerald AI, received the signal about grid conditions and adjusted the data center's flexible computing workloads. Work that could wait was slowed or rescheduled, while higher-priority services continued operating. The goal was to reduce electricity demand when the grid was constrained without interrupting critical AI workloads — exactly what Silicon Valley Power needed. When the reduction showed on screen, everyone cheered. Sivaram said, "We were watching with bated breath. It was our first time deploying across thousands of NVIDIA GPUs." His head of product, Mansi Shah, was emotional. "This feels kind of like a SpaceX rocket launch," she said.
Silicon Valley Power has since sent more than 200 demand signals to that AI factory, and it worked every single time. When the signal arrived, the factory's power dropped from four megawatts to three, automatically, while every high-priority job kept running.
This points to a much bigger opportunity: a path to unlocking the power America's AI factories need without waiting a decade to build new transmission lines. At the AI Infra Summit, Ian Buck, NVIDIA's vice president of hyperscale and high-performance computing, made AI factory efficiency the centerpiece of his infrastructure keynote. Results from cloud provider Lambda's first validation, released the same day, put numbers to it: a fixed power budget can support 24 percent more token throughput when managed intelligently. Dave Ward, president of cloud services at Lambda, said, "With our proof of concept, we believe we've moved beyond the limitation of fixed power budgets. NVIDIA DSX MaxLPS paves the way to reclaiming stranded capacity and converting it into real-world usage, with significantly more compute density in the same footprint."
When Silicon Valley Power called that August evening, Conductor executed against a predefined workload hierarchy: lowest-priority jobs yielded, high-priority inference kept running, and power fell from four megawatts to three. Automated. No operator required.
In the AI factory economy, power is the constraint. Work per gigawatt is the metric. NVIDIA DSX extends efficiency discipline into the AI workload itself. Smarter rack provisioning puts power where workloads actually need it. Operational intelligence — tighter scheduling, faster restarts, leaner checkpointing — keeps GPUs running rather than waiting. Jensen Huang, NVIDIA founder and CEO, has said, "A one-gigawatt factory will never become a two-gigawatt factory."
Lambda ran NVIDIA's DSX MaxLPS software on a five-rack, 19-node cluster. By running 19 nodes within the same power budget as 16 nodes at full power, Lambda achieved 24 percent more cluster-wide token throughput — from roughly 4 million tokens per second to 5 million. Performance per watt improved by 23 percent. NVIDIA projects DSX MaxLPS can enable up to 40 percent more GPU capacity for next-generation Vera Rubin NVL72 AI factories within the same megawatt power budget in suitable deployment environments.
NVIDIA's Eos AI factory is running Emerald AI Conductor as a participant in Silicon Valley Power's Flexible Load Interconnect Program, the first commercial grid utility program designed to treat AI factories as dispatchable resources. When Silicon Valley Power sends a signal, Conductor responds in under a minute. The first dedicated DSX Flex commercial deployment will be the Manassas, Virginia facility: a 96-megawatt Vera Rubin AI factory at NVIDIA's AI Factory Research Center.
NVIDIA's 800V DC power architecture is designed to reduce conversion complexity, improve power delivery efficiency and support denser accelerated computing racks. DSX is incorporating 800V DC into its reference designs.
No single component can optimize an AI factory on its own. A faster GPU still waits on the network. Power can be stranded by bad provisioning. The only reliable path to more tokens per megawatt is to optimize the whole factory using DSX Sim, DSX OS, DSX Exchange, and DSX Reference Designs. The fundamental question remains: how much useful work does the factory produce per megawatt consumed? With NVIDIA DSX, that is the baseline for what an AI factory is supposed to do.