NVIDIA is advancing the next era of AI inference by integrating its Vera Rubin NVL72 platform with the new Groq 3 LPX ac
The next phase of artificial intelligence inference will be determined by how every layer of the AI factory operates as a unified system rather than relying on isolated hardware breakthroughs. To address this, NVIDIA is extending the Vera Rubin NVL72 rack-scale system with the newly announced Groq 3 LPX accelerator, specifically engineered for rapid token generation in agentic AI applications. The system is now in full production and achieved remarkable results in an Artificial Analysis benchmark running the open-source Gemma 4 31B model, delivering 3,400 output tokens per second for 100,000-token long-context scenarios that are essential for agentic systems, outperforming the closest competing platform by four times.
Global industry partners are already integrating the Vera Rubin platform solutions into their operations. Nebius has become the first AI cloud to adopt the Groq 3 LPX, enabling developers to build highly responsive agentic applications, coding assistants, and real-time AI experiences at scale within its Token Factory environment. CoreWeave has moved Spectrum-X Multiplane into production, utilizing multiple parallel switches to interconnect Vera Rubin racks and establish high-bandwidth, flat, and lossless AI networks. Meanwhile, SpaceXAI announced that NVIDIA Vera CPUs will power its upcoming agentic AI architecture, handling CPU-intensive tasks such as orchestration, tool execution, code processing, data management, and simulation across both terrestrial data centers and orbital satellites.
As artificial intelligence transitions from model training to reasoning and autonomous agent collaboration, inference has emerged as the primary frontier. These advanced workloads require infrastructure optimized for extreme throughput, minimal latency, and economic efficiency at massive scale. At the Hot Chips conference in Palo Alto, California, NVIDIA highlighted how extreme codesign is transforming the entire AI pipeline. By architecting compute, networking, and inference acceleration as a cohesive unit, the company enables customers to deploy purpose-built infrastructure for long-context processing and multi-agent workflows. The Groq 3 LPX introduces a specialized low-latency architecture that works alongside the Vera Rubin NVL72, addressing the critical challenge of decode latency where tiny delays compound across complex agent chains. While Rubin GPUs manage large-scale context processing, the LPX units accelerate latency-sensitive decoding, eliminating the traditional tradeoff between speed and throughput.
Unlike standalone accelerators, the Groq 3 LPX leverages extreme codesign to combine GPU and LPU strengths, allowing both to jointly compute every layer of an AI model. At scale, fleets of these processors function as a single deterministic inference engine, with a single rack deployment capable of housing 256 LP30 accelerators linked through direct chip-to-chip connections. Network performance remains equally vital as AI factories expand, which is why NVIDIA is showcasing Spectrum-X Multiplane at Hot Chips. This latest iteration of the hardware-accelerated Spectrum-X Ethernet architecture scales to unprecedented sizes without introducing additional network tiers, thereby avoiding extra latency, jitter, and costs. The end-to-end Ethernet platform, comprising switches, SuperNICs, and specialized software, delivers 1.6 times better AI networking performance compared to standard off-the-shelf Ethernet while maintaining consistent results in multi-tenant environments. Together, these innovations position NVIDIA to transform AI factories into integrated engines capable of converting growing token volumes into reliable revenue, fulfilling what the company describes as a “token factory” vision, with further optimizations and performance gains expected as the ecosystem evolves.