Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

NVIDIA confirms full-scale production of the Groq 3 LPX accelerator while highlighting Vera Rubin’s multi-fold inference efficiency gains.

Mass production accelerates agentic AI deployment cycles and validates next-generation inference architecture ahead of hyperscaler demand.
Trade pressSlicast · August 25, 2026 · US · Source: Google News
importance 92

NVIDIA has announced the full-scale production launch of Groq 3 LPX for interactive AI inference, positioning it as a key component of the Vera Rubin platform. Designed to emphasize ultra-low-latency token generation, Groq 3 LPX does not aim to replace GPUs with specialized inference chips; rather, it complements the Vera Rubin architecture by enhancing traditional GPU capabilities in low-latency token delivery. On Monday, the 24th (Eastern Time), NVIDIA revealed that in Artificial Analysis tests, Groq 3 LPX achieved an output rate of 3,400 tokens per second using the Gemma 4 31B model with a 100,000-token context window, setting a new benchmark for that model. For low-latency workloads such as agentic programming, its response speed was up to four times faster than recent competing platforms.

Concurrently, NVIDIA released new performance data for the Vera Rubin NVL72 under real-world agentic workloads. Leveraging SemiAnalysis’s AgentX workloads and the DeepSeek V4 Pro model, the Vera Rubin platform delivered a throughput per megawatt up to 30 times higher than the previous-generation GB300 NVL72, while reducing the cost per token by up to 35 times. NVIDIA emphasized that agentic tasks fundamentally differ from traditional chat or summarization workflows. Because task contexts accumulate continuously over hundreds of steps—potentially reaching hundreds of thousands of tokens—the speed of token generation at each step directly dictates the total time required to complete a task. Agents routinely read files, invoke tools, execute code, verify outputs, and iterate through hundreds or even thousands of reasoning cycles. Consequently, Agentic AI presents two distinct computational challenges: efficiently processing increasingly large-scale contexts, and continuously generating tokens with ultra-low latency. The Vera Rubin NVL72 addresses broader training, inference, and context processing, while Groq 3 LPX is specifically optimized for single-user interaction speed to maximize token generation rates.

The shift toward these metrics marks a strategic evolution in AI infrastructure competition. Moving beyond traditional model training and inference benchmarks, NVIDIA is refocusing the industry on token generation speed, energy consumption, and cost efficiency for Agentic AI. While single-dimensional metrics like 3,400 tokens per second remain notable, NVIDIA highlighted throughput per megawatt and cost per token as more critical indicators of real-world value. The reported 30-fold increase in throughput per megawatt reflects a significant improvement in overall AI inference efficiency rather than a straightforward 30x boost in raw chip speed. This distinction is vital given that traditional large model services prioritize latency and throughput per individual request, whereas agent workloads continuously consume tokens. Citing data from OpenRouter, NVIDIA noted that AI agent workloads can consume up to 15 times more tokens than simple chat requests. An agent may query databases, search documents, invoke sub-agents for analysis, execute code, and repeatedly verify results, with context expanding throughout the entire process. As power and data center resources increasingly constrain AI expansion, maximizing tokens generated per unit of electricity has become more impactful than merely boosting peak computing power.

On the commercial front, AI cloud service provider Nebius will be the first vendor to adopt Groq 3 LPX, planning to deploy it on the Nebius Token Factory production inference platform. Groq itself also intends to be among the earliest adopters. This rollout signals that the capabilities NVIDIA acquired through its prior purchase of Groq technology are now being fully integrated into the Vera Rubin data center product portfolio.

Expanding beyond terrestrial deployments, NVIDIA announced that SpaceXAI will adopt the Vera CPU to accelerate next-generation Agentic AI applications and scale its Vera Rubin-based computing infrastructure to several gigawatts. Additionally, SpaceXAI plans to deploy an optimized version of the Vera Rubin NVL72 in orbit for use in its first-generation Starmind AI satellites. From an architectural perspective,

Read the original
NVIDIA confirms full-scale production of the… · Slicast