Friday, September 11, 2026
AI 인프라 · 뉴스 & 분석
반도체·하드웨어리포트
반도체·하드웨어 · 리포트

NVIDIA가 Groq 3 LPX 추론 가속기를 Vera Rubin 플랫폼과 함께 본격 생산에 진입하여 전달

NVIDIA 공식 — 로드맵/제품 직접 확인
공식 공시Slicast · September 11, 2026 · 미국 · 출처: NVIDIA Blog

NVIDIA is reshaping its AI infrastructure strategy around what it calls the "AI factory," a unified system where compute, networking and inference work together as an integrated whole rather than optimized components operating in isolation. The company announced that the Vera Rubin NVL72 rack-scale system is being extended with the NVIDIA Groq 3 LPX, a specialized inference accelerator purpose-built for token generation in agentic systems.

According to an Artificial Analysis benchmark running Gemma 4 31B, an open source agentic model, the Groq 3 LPX delivered 3,400 output tokens per second for 100,000-token long-context use cases, achieving 4x faster performance than the nearest alternative platform. The accelerator is now in full production.

Industry adoption is already underway. Nebius became the first AI cloud to adopt the Groq 3 LPX, integrating it with Vera Rubin in its Token Factory offering to help developers build highly interactive agents and real-time AI experiences at scale. CoreWeave is deploying Spectrum-X Multiplane in production, which connects multiple Vera Rubin racks through parallel switches to provide high-bandwidth, flat and lossless AI networks. SpaceXAI announced plans to build its future AI architecture around NVIDIA Vera Rubin, starting with NVIDIA Vera CPUs to handle CPU-intensive agentic work including orchestration, tool use, code execution, data processing and simulation.

The shift in AI infrastructure reflects a fundamental change in market demands. As AI moves from training to reasoning and agentic applications, inference has become the new frontier. Agentic systems generate more tokens, process dramatically larger context windows and increasingly collaborate with other AI systems to solve complex problems. Decode latency has emerged as a new performance challenge — as AI agents reason, use tools and interact with other systems, they generate responses one token at a time, and tiny delays multiply across complex chains of work.

NVIDIA Groq 3 LPX was codesigned to work alongside Vera Rubin NVL72 such that Rubin GPUs handle large-scale context processing while LPX accelerates latency-sensitive decode workloads. The result is designed to be faster, more predictable token generation that helps AI factories deliver responsive reasoning and greater infrastructure efficiency. At scale, fleets of LPUs can be connected through direct chip-to-chip links, with a rack-scale Groq 3 LPX deployment capable of including 256 LP30 accelerators, creating what NVIDIA describes as a highly efficient inference engine built for modern AI factories.

원문 보기