Benchmarks show the NVIDIA Groq 3 LPX accelerator delivering 3,431 transactions per second when deployed on the Vera Rubin platform.
NVIDIA’s Groq 3 LPX has established a new performance standard for AI inference, delivering 3,431 tokens per second (TPS) on a 100K context benchmark. According to an official blog post released on August 24, 2026, this achievement redefines capabilities for high-interactivity and long-context AI workloads. The benchmark was conducted by Artificial Analysis using the Gemma 4 31B model, demonstrating the system’s capacity to handle demanding tasks on NVIDIA’s advanced Vera Rubin platform.
Built around NVIDIA’s LP30 accelerators and a rack-scale architecture, the Groq 3 LPX delivers 315 PFLOPS of FP8 inference compute alongside 128 GB of SRAM. Its design prioritizes deterministic execution, low-latency token generation, and fine-grained scheduling, enabling exceptional performance in multiturn agentic sessions and interactive workflows where context expands significantly with each user interaction. Traditional models frequently struggle to maintain responsiveness at this scale, particularly when processing lengthy input contexts.
During testing, Artificial Analysis evaluated the system using a 100K input context length to measure its output generation speed. This capability is critical for agentic AI applications, such as coding or complex reasoning workflows, which require processing large volumes of accumulated context. NVIDIA reports that these speeds enable the generation of 5,000 tokens in just 1.5 seconds, compared to 50 seconds at 100 TPS—an order-of-magnitude improvement in efficiency. On a 10K context benchmark, the system achieved 3,382 TPS with minimal latency variation. Additionally, NVIDIA’s SPEED-Bench tests for coding tasks recorded a median throughput of 4,767 output tokens per second, with 20% of tasks exceeding 5,500 TPS, underscoring the LPX’s versatility across diverse use cases.
Positioned as a foundational component within NVIDIA’s Vera Rubin platform, the Groq 3 LPX is designed for AI factories requiring high-interactivity serving tiers. Its ability to manage over 100K context tokens while sustaining low latency could transform industries dependent on complex, multiturn AI interactions, including customer service automation, large-scale coding assistants, and real-time decision-making systems. A key driver of this performance is the LPX’s compiler-scheduled workload planning, which minimizes communication overhead across its 256 interconnected LPUs. By overlapping computation and communication at a fine-grained level, the architecture achieves unmatched efficiency, even at small batch sizes where conventional tensor parallelism techniques typically underperform.
This announcement arrives at a pivotal moment for NVIDIA, which has solidified its leadership in AI hardware. While the Groq 3 LPX is not a cryptocurrency-focused product, its implications for AI-driven sectors are substantial. Enhanced by the Groq 3 LPX, the Vera Rubin platform positions NVIDIA to dominate high-demand AI workloads, ranging from real-time inferencing to generative AI applications. In financial markets, NVIDIA (NVDA) shares recently traded at $210.18, reflecting a 2.11% decline over the previous 24 hours amid broader market softness. Nevertheless, advancements in AI inference technology continue to support the company’s long-term growth trajectory, particularly within high-margin enterprise solutions. Separately, NVIDIA has officially denied recent rumors regarding a China-specific LPU product, clarifying that no such roadmap exists—a development that may alleviate geopolitical concerns among investors.
The demonstrated performance of the Groq 3 LPX paves the way for new AI applications that demand both extreme speed and expansive context handling. NVIDIA has indicated plans to sustain these throughput levels even with multi-hundred-thousand token contexts, potentially further revolutionizing agentic AI use cases. Given its robust performance metrics and strategic integration with the Vera Rubin architecture, the Groq 3 LPX is poised to set the benchmark for high-interactivity AI systems in the coming years.