Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeHeadlinesReport
Headlines · Report

NVIDIA has begun full-scale mass production of the Groq 3 LPX, with initial deployments confirmed at Nebius.

Early cloud adoption by Nebius validates low-latency inference demand and sets a benchmark for next-gen AI workload performance.
Trade pressSlicast · August 25, 2026 · US · Source: Google News
importance 94

On Monday, NVIDIA (NVDA.US) announced that its Groq 3 LPX rack-level systems have entered full-scale mass production, marking the formal commercialization of its low-latency AI inference technology. This milestone follows the company’s acquisition of Groq-related assets in December last year for approximately $20 billion, representing the largest transaction in NVIDIA’s history. The first batch of systems will be deployed with AI cloud computing provider NEBIUS (NBIS.US) and is scheduled to go live later this year.

Dion Harris, Senior Director at NVIDIA, told the media that the Groq 3 LPX will be deployed alongside NVIDIA Vera central processing units and Rubin graphics processing units in Nebius data centers. A defining feature of the Groq chip architecture is the integration of 500MB of high-speed SRAM directly on the chip, which mitigates data transmission bottlenecks associated with traditional memory access and accelerates response times during AI model inference. While NVIDIA’s primary GPUs are manufactured by Taiwan Semiconductor (TSM.US), Groq chips are produced by Samsung. Each LPX rack currently integrates 256 Groq 3 chips.

According to benchmark data from Artificial Analysis cited by NVIDIA, the Groq 3 LPX system can generate approximately 3,400 tokens per second. In generative AI, a token serves as the fundamental unit for text processing and generation; higher token throughput enables faster responses, which is critical for latency-sensitive applications like AI programming and real-time agents. Harris noted that lower latency allows cloud providers to offer premium, higher-priced services to clients with stringent response requirements. However, he emphasized that Groq chips are not intended to replace traditional GPUs. “This is not about replacing GPUs, but rather using the most cost-effective and high-performance processors for different parts of the workload,” Harris stated. GPUs will continue to serve as the versatile backbone for both model training and general inference, while Groq focuses specifically on the “decoding” phase of content generation.

This approach signals NVIDIA’s push toward a segmented AI computing architecture: Vera CPUs and Rubin GPUs will manage broad computational workloads, while Groq chips are optimized for latency-critical inference tasks. As the industry pivots from large-scale model training to rapidly expanding inference demand, competition in the low-latency segment is intensifying. Earlier this year, AMD (AMD.US) announced the integration of rack-scale AI systems with Cerebras (CBRS.US) chips to target the same niche. The sector’s growing importance was highlighted by OpenAI’s recently announced Ultrafast mode, which promises 750 tokens per second powered by Cerebras infrastructure. NVIDIA cautioned that its 3,400-token benchmark should not be viewed as a direct performance comparison due to varying test conditions and application scenarios.

Concurrently, NVIDIA is accelerating shipments of the Vera Rubin system, which entered production earlier this year. When unveiling the platform in March this year, CEO Jensen Huang projected that cumulative sales spanning from current Blackwell chips to next-generation Vera Rubin systems could reach $1 trillion by 2027. Huang also revealed plans to dedicate roughly one-quarter of data center capacity for AI programming applications to Groq chips, with the remainder allocated to the Vera Rubin system. This allocation underscores NVIDIA’s strategic emphasis on the low-latency inference market, where processing speed is becoming a decisive competitive advantage for cloud firms and AI developers. The official ramp-up of Groq 3 LPX production marks NVIDIA’s evolution from a GPU-centric supplier to a provider of specialized, workload-tailored computing architectures. Investors will closely monitor these developments when the company reports earnings this Wednesday, alongside demand trends for Blackwell and Vera Rubin.

Read the original
NVIDIA has begun full-scale mass production of… · Slicast