NVIDIA argues that power efficiency—measured in performance per watt—is the defining constraint for AI factories, and po
Power is AI infrastructure's inescapable constraint. The number of tokens an AI factory can generate within a fixed power budget determines its revenue and profitability. Performance per watt, a metric that cannot be gamed but only earned through real-world results, is the foundation for AI factories. As agentic AI drives token demand higher, the infrastructure decisions organizations make today will determine who scales and who doesn't in a power-constrained world.
Virtually every frontier AI model today runs on a mixture-of-experts architecture. Serving these large-scale models efficiently means GPU domain size—the number of GPUs connected over an ultrafast, scale-up interconnect—matters, and bigger is better. While the NVIDIA Hopper generation set the standard with an eight-GPU domain, the scale of frontier AI today has outgrown it. Serving MoE with a 72-GPU domain demands full-stack codesign and the operational depth earned from running these models under real production load.
With the NVIDIA Blackwell NVL72 platform, that foundation is already built and proven, delivering the highest performance per watt to maximize revenues and the lowest token cost to maximize profit margins. The NVIDIA Vera Rubin platform builds upon this foundation next to further elevate rack-scale energy efficiency. Across the newest generation of leading open models, NVIDIA GB300 NVL72 delivers up to 25x performance per watt compared with the NVIDIA Hopper generation, showcasing that MoE performance improves when moving from an 8-GPU to 72-GPU domain size.
Different workloads demand different operating points: some optimize for latency, others for throughput and cost, and most need to move between the two. Rather than showcasing a single number, NVIDIA presents Pareto curves for each model and provides tools such as DynoSim to help teams find their optimal point on the Pareto frontier before spending GPU-hours on validation.
The performance per watt NVIDIA Blackwell delivers results from extreme codesign where every component of the rack-scale system, from silicon to software, is designed together to maximize token throughput for AI inference workloads. The NVIDIA NVLink Switch, now in its sixth generation with the Vera Rubin platform, is purpose-built to unlock massive scale-up GPU domains and includes capabilities designed specifically for AI workloads such as SHARP, which performs in-network computing directly in the switch. NVIDIA's inference software stack, including NVIDIA Dynamo, TensorRT LLM, SGLang and vLLM, runs the full range of optimizations from NVFP4 quantization to KV cache offloading. On DeepSeek V4, performance per watt improved by up to 5x in a single month through software improvements alone.
In AI factories, power lost to cooling and rack-level inefficiencies means only about 60% of electricity pulled from the grid turns into useful AI work. NVIDIA DSX MaxLPS, the power-and-efficiency software in the NVIDIA DSX platform, closes that gap by shifting power between GPUs and racks in real time and using techniques like power steering to wring more performance. This enables operators to run up to 40% more GPUs within the same power budget.
Rack-scale reliability at AI factory scale is hard-won, introducing failure modes that single-node deployments never encounter. NVIDIA Blackwell NVL72 systems continue to set the standard across diverse models and production use cases. Leading AI labs including Anthropic, OpenAI and SpaceXAI use NVIDIA Blackwell NVL72 systems to run inference, along with inference service providers and AI natives deploying open models in production. CoreWeave has deployed Kimi K2.6 on NVIDIA GB300 NVL72 combining NVFP4 quantization and EAGLE3 speculative decoding. Perplexity runs Qwen3 235B and post-trained Qwen3.5-397B-A17B on NVIDIA GB200 NVL72 for its AI agent platform, serving millions of queries daily. Fireworks AI deploys GLM 5.2 on the NVIDIA Blackwell platform for customers including Cursor and Factory AI. This accumulated production experience across frontier models and real-world deployments gives NVIDIA Vera Rubin its foundation.