Friday, August 28, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Nvidia has reached full production volume for its low-latency AI inference LPX rack systems.

Scaling LPX production accelerates the deployment of optimized inference clusters, reducing time-to-market for generative AI applications and increasing utilization rates for existing GPU fleets.
Trade pressSlicast · August 27, 2026 · US · Source: Google News
importance 79

Nvidia has confirmed that its LPX racks for AI inference accelerators have entered full production. Unveiled at GTC in March, the rack-scale platform was developed following Nvidia’s acqui-hire of Groq. The liquid-cooled LPX system houses 256 Groq 3 language processing units (LPUs), interconnected via 640 Tb/s of scale-up bandwidth. Within each rack, BlueField-4 data processing units (DPUs), Vera central processing unit (CPU) racks, and STX storage servers are integrated using the recently introduced Spectrum-6 Ethernet networking technology.

Rather than replacing Nvidia’s flagship NVL72 platform, LPX serves as a complementary solution tailored for operators requiring low-latency AI inference capabilities. During this week’s Hot Chips conference in Palo Alto, Nvidia presented industry benchmarking results demonstrating that the LPX platform can sustain 3,400 output tokens per second. The company states that the architecture can reduce the execution time for agentic workflows from hours to minutes, delivering four-times faster responsiveness for AI agents and other latency-sensitive applications.

“The LPX will transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency, and responsiveness,” said Nvidia founder and CEO Jensen Huang. “Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model (LLM) inference with an unprecedented leap in performance and efficiency. Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

The LPX platform is scheduled for availability later this year. Among its early adopters is Nebius, which plans to integrate the Groq 3 LPX architecture into its Token Factory offering. Nebius CTO Danila Shtan stated that the deployment will make “every step of an agent’s loop feel instant.”

Read the original
Nvidia has reached full production volume for… · Slicast