Nvidia confirms Groq 3 LPX has entered full production with deployment racks expected online this year.
Nvidia has moved its Groq 3 LPX racks into full production, expanding its strategic push into low-latency artificial intelligence inference alongside its existing GPU and CPU offerings as demand for faster AI agents continues to grow.
Announced on Monday, the commercialization of the Groq 3 LPX marks the market debut of technology derived from Nvidia’s largest acquisition on record. The racks will be deployed alongside Vera central processors and Rubin graphics processors at neocloud provider Nebius, with initial deployments expected later this year.
Nvidia’s decision to manufacture and distribute Groq chips underscores the increasing industry priority placed on low-latency inference. This capability is essential for ensuring AI agents remain highly responsive without perceptible delays, particularly in coding applications. Cloud providers can also command higher pricing for these high-speed inference workloads.
In December, Nvidia acquired assets from chip startup Groq for $20 billion, making it one of the company’s largest acquisitions. Groq’s architecture includes 500 megabytes of high-speed SRAM directly on the chip, a design intended to reduce memory-related bottlenecks. Groq chips are manufactured by Samsung, while Taiwan Semiconductor Manufacturing Company (TSMC) manufactures Nvidia’s GPUs.
Each LPX rack integrates 256 Groq 3 chips. Based on a benchmark from Artificial Analysis, Nvidia claims a Groq 3 LPX rack can deliver up to 3,400 tokens per second.
This development positions Nvidia in direct competition with other firms targeting low-latency AI inference. Earlier this year, AMD announced it would integrate its rack-scale systems with chips from Cerebras. Meanwhile, OpenAI recently introduced an Ultrafast mode powered by Cerebras, which promises speeds of up to 750 tokens per second.
Low-latency chips such as Groq are not designed to replace GPUs, which remain the primary workhorses for AI workloads, including both training and inference. Instead, Groq chips primarily focus on the “decode” phase of serving AI models.
Concurrently, Nvidia is ramping up shipments of its Vera Rubin systems, which entered production earlier this year. At the March unveiling of the Vera Rubin platform and Groq 3 LPX, Nvidia CEO Jensen Huang projected $1 trillion in cumulative sales from Blackwell and Vera Rubin systems through 2027.