Nvidia has begun production of Groq AI racks following its $20 billion acquisition of the inference chipmaker.
Nvidia announced Monday that its Groq 3 LPX artificial intelligence racks have entered full production, commercializing technology acquired through the chipmaker’s record $20 billion purchase of Groq assets. The acquisition, finalized in December, stands as Nvidia’s largest-ever transaction.
Designed specifically for low-latency inference—the process by which trained AI models generate responses—the Groq 3 LPX systems address a critical bottleneck in AI performance. Faster inference is especially vital for AI agents and coding assistants, where even minor delays degrade user experience. Each rack integrates 256 Groq 3 chips and delivers approximately 3,400 tokens per second, according to a benchmark cited by Nvidia. To minimize memory-related bottlenecks, the Groq architecture embeds 500 megabytes of high-speed static random-access memory directly onto each chip. The chips themselves are manufactured by Samsung Electronics, while Taiwan Semiconductor Manufacturing Company continues to produce Nvidia’s graphics processors.
Nvidia emphasized that these specialized systems are intended to complement, rather than replace, traditional GPUs. While GPUs handle both AI model training and inference, Groq chips are optimized for the latency-sensitive “decode” phase of model execution. The new racks will be deployed alongside Nvidia’s Vera central processing units and Rubin graphics processing units at cloud infrastructure provider Nebius, with operations expected to begin later this year. Concurrently, Nvidia is ramping up shipments of its Vera Rubin systems, which entered production earlier this year.
Looking ahead, Nvidia CEO Jensen Huang stated in March that the company projects cumulative sales from its Blackwell and Vera Rubin platforms to reach $1 trillion through 2027. Huang also noted that a quarter of all data center capacity dedicated to coding applications will utilize Groq chips.
The push into specialized inference hardware reflects intensifying competition as tech firms race to make AI services faster and more cost-effective. In response, Nvidia rival Advanced Micro Devices has announced plans to integrate rack-scale systems with inference chips developed by Cerebras.