Friday, October 2, 2026
AI 인프라 · 뉴스 & 분석
홈 › 컴퓨트·클라우드 › 리포트
컴퓨트·클라우드 · 리포트

CoreWeave가 Nvidia의 Vera Rubin NVL72 GPU 클러스터를 상용 규모로 제공 중이며, Cognition이 초기 고객입니다.

주요 neocloud가 Nvidia의 차세대 GPU 가용성을 직접 채널을 넘어 확대하며, AI 워크로드용 Vera Rubin 클러스터 도입을 가속화합니다.
업계 전문지Slicast · 2026년 9월 30일 13:00 UTC · 글로벌 · 출처: Data Center Dynamics
중요도 85

CoreWeave is now offering Nvidia's Vera Rubin NVL72 system on its cloud platform. The announcement came at CoreWeave's Fully Connected conference in San Francisco, just two months after the company began operating what it claims was the first fully functional Vera Rubin NVL72 rack in production.

Cognition is the first customer to adopt the platform, deploying Vera Rubin systems in early September. According to Silas Alberti, SVP research & founding team at Cognition, the company has achieved up to a 4.8x increase in total token throughput for its SWE-2 inference workloads on the new generation.

"Bringing up Nvidia Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations," said Chen Goldberg, executive vice president of product & engineering at CoreWeave. "With customers like Cognition, that investment shows up in the ability to get production workloads running within days. When it comes to agentic tasks, long contexts, repeated model calls, and thousands of concurrent tasks put pressure on the entire platform. Our job is to make compute, networking, and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves."

The Rubin platform comprises six chips: the NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet Switch, alongside the Vera CPU and Rubin GPU. Vera succeeds Nvidia's Grace CPU, while Rubin will succeed Blackwell GPUs. Nvidia claims Rubin will deliver 5x inference performance and 3.5x training performance compared to Blackwell.

The NVL72 system—Nvidia's rack-scale offering with 36 CPUs and 72 GPUs—is 100 percent liquid-cooled with cable-free modular tray designs. Nvidia claims this design reduces installation time from two hours to five minutes.

CoreWeave is also planning to offer the Vera CPU as a standalone product. According to Corey Sanders, SVP of product at CoreWeave, the company is in early stages but expects customers to begin testing in the coming weeks, with a few already having early access. Deployed at rack-scale, CoreWeave will offer 128 Vera CPUs and 11,264 cores per rack, with BlueField-4 DPUs and Spectrum-X Ethernet switching providing secure, high-performance connectivity between Vera nodes. The CPU will run on CoreWeave as a bare-metal offering, designed for AI agents and new workloads.

"General-purpose infrastructure bottlenecks agentic AI; Vera is the first CPU explicitly designed to accelerate it," said Goldberg. "Our platform natively enables Vera with products like CoreWeave Sandboxes out of the box. Teams can instantly spin up thousands of isolated environments, removing operational friction and accelerating the entire AI loop on day one."

CoreWeave also announced a new customer contract with Ennoble Care, a home-based primary, palliative, and hospice care company. Ennoble is deploying Nvidia RTX Pro 6000 Blackwell Server Edition nodes on CoreWeave Kubernetes Service to support AI in clinical workloads including summarization, documentation, and clinical decision support.

"We've built a full-stack, ONC-certified EMR purpose-built for home-based primary care, and we're now developing multiple AI agents on top of it, both to augment clinical delivery and to automate back-office functions," said Jonathan Taylor, CTO of Ennoble Care. "CoreWeave is the right solution to help maintain our clinical standards. It gives us reserved capacity we can count on, a Kubernetes environment our team can move into quickly, and engineers who answer the phone. That combination allows us to scale the AI capabilities already built into our care delivery workflows and proprietary electronic medical record system, extending them to tens of thousands of additional patients as we continue to grow."

"Clinical inference is one of the most demanding places AI can run. The workload is continuous, the latency budget is short, and the compliance requirements are absolute," said Jon Jones, chief revenue officer at CoreWeave. "That is what The Essential Cloud for AI means in practice. It's a cloud ready for the workloads that matter most, in the industries where precision is everything."

CoreWeave also launched CoreWeave Forge, a development layer that enables teams to build, improve, and evaluate AI models and agents in a single environment. The offering is now generally available.

원문 보기
CoreWeave가 Nvidia의 Vera Rubin NVL72 GPU 클러스터를… · Slicast