Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeCapital MarketsReport
Capital Markets · Report

NVIDIA confirms that the AI inference accelerator developed under its $20 billion strategic partnership with Groq has officially entered full production.

The massive capital commitment validates long-term custom silicon strategies, locking in dedicated inference capacity ahead of broader market saturation.
Trade pressSlicast · August 25, 2026 · US · Source: Google News
importance 90

Nvidia ($NVDA) announced Monday that its Groq 3 LPX inference chip has entered full production, marking a key milestone for the technology underlying the company’s $20 billion acquisition of chip startup Groq’s assets in December—the largest transaction Nvidia has ever completed.

The Groq 3 LPX extends Nvidia’s Vera Rubin platform and is engineered to accelerate the “decode” phase of AI inference, which dictates token generation speed for end users. Each LPX rack houses 256 individual Groq 3 chips. Fabricated by Samsung, the chip integrates 500 megabytes of SRAM directly onto the die, bypassing the memory bandwidth bottlenecks that typically constrain competing inference accelerators, according to CNBC.

In benchmarking conducted by Artificial Analysis, the Groq 3 LPX achieved 3,400 output tokens per second while running the open-source Gemma 4 31B model across a 100,000-token context window. Nvidia states this performance is four times faster than the closest competing platform for latency-sensitive workloads.

AI cloud provider Nebius will be the first to deploy the chip through its Nebius Token Factory service. According to CNBC, these racks are scheduled to go live later this year alongside Vera central processors and Rubin graphics processors. Following Nebius, the originally independent Groq startup plans to become one of the earliest adopters of the technology.

“Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate,” said Danila Shtan, chief technology officer at Nebius.

Nvidia senior director Dion Harris emphasized the chip’s commercial value in enabling cloud providers to launch premium service tiers. “For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive” service agreements, Harris noted. He clarified that the Groq 3 LPX is designed to complement, rather than replace, Nvidia’s existing GPU portfolio, which remains critical for model training and general-purpose inference. “This isn't about replacing GPUs,” Harris said. “It's about using the right price, right processor for the right part of the workload.”

Shipments of Vera Rubin systems are now accelerating, following the start of production earlier in 2026. Nvidia is set to report its next quarterly earnings on Wednesday.

Read the original
NVIDIA confirms that the AI inference… · Slicast