CoreWeave는 Livingston 시설에서 다중 랙 NVIDIA Vera Rubin NVL72 클러스터를 성공적으로 가동했습니다.
LIVINGSTON, N.J., Sept. 16, 2026 — CoreWeave, Inc. today announced the bring-up of multi-rack NVIDIA Vera Rubin NVL72 systems on CoreWeave Cloud, deploying hundreds of NVIDIA Rubin GPUs within a single scale-out cluster optimized for agentic AI. The company also unveiled two new capabilities for CoreWeave AI Object Storage: cross-region write acceleration and a new Archive tier. These features keep the data critical to these workloads close to the GPUs, ensuring continuous productivity and accelerating the AI development loop.
Multi-rack Vera Rubin NVL72 clusters enable training and inference jobs to span hundreds of Rubin GPUs simultaneously. This architecture is particularly vital for multi-step agentic workloads, which are highly sensitive to data-access latency; delays can compound across repeated model calls and tool interactions. Cross-region write acceleration eliminates wait times even when operating across multiple geographic regions. By writing data locally while CoreWeave replicates it to another region in the background, an agent’s intermediate state, retrieved context, and outputs move at the speed of its reasoning. The GPUs do not stall, and neither does the development loop.
“CoreWeave was the first AI cloud provider to validate and bring up a Vera Rubin NVL72, demonstrating that this advanced rack-scale architecture could operate as a reliable, high-performance cloud service,” said Chen Goldberg, executive vice president of product & engineering at CoreWeave. “With multi-rack Vera Rubin, we are connecting hundreds of Rubin GPUs as a single scale-out cluster. For customers building agentic AI, that means greater scale, faster iteration, and higher productivity as models and agents continuously learn and improve.”
Scaling agentic AI with multi-rack NVIDIA Vera Rubin NVL72 requires unifying hardware into a cohesive system. A single NVIDIA Vera Rubin NVL72 rack pairs 72 Rubin GPUs with 36 Vera CPUs, NVIDIA NVLink 6, ConnectX-9 SuperNICs, and BlueField-4 DPUs. Through multi-rack deployments, CoreWeave unifies racks containing hundreds of accelerators using NVIDIA Spectrum-X Ethernet networking into a single scale-out cluster. This configuration delivers the capacity to train larger models, serve more demanding inference workloads, and run reinforcement learning at scale. Achieving this requires precise engineering across every layer—including compute, networking, storage, cooling, power, firmware, and software—to operate as one coordinated system.
CoreWeave brings multiple racks online as a unified system through three key approaches: automating rack lifecycle control, validating performance from components to systems, and scaling the network alongside the GPUs. Racks arrive as hardware requiring connection and validation, but CoreWeave Mission Control automates setup via the Rack LifeCycle Controller, coordinating hardware detection, firmware updates, validation, power, and cooling. Racky handles rack control, while Valvey executes cooling actions. Performance validation combines NVIDIA field diagnostics with full-rack workload testing, comparing every result against established baselines. Leveraging years of real-world experience, the company ensures that only racks meeting strict system-level performance thresholds enter production, guaranteeing optimal GPU performance. Network scaling is handled by equipping each Rubin GPU with two NVIDIA Connect X-9 SuperNICs, delivering 1.6 Tb/s of connectivity per GPU across multiplane, multirail paths. This supports approximately 128,000 GPUs per rail in a non-blocking fabric, allowing additional racks to be integrated without redesigning the underlying fabric.
AI performance hinges on moving data as quickly as compute scales. CoreWeave AI Object Storage Local Object Transport Accelerator (LOTA) brings data closer to AI workloads through managed caching on each CoreWeave Kubernetes Service node. This delivers read speeds matching local NVMe drives and reduces latency by 8x compared to traditional storage clusters. LOTA provides up to 7 GB/s of throughput per GPU and scales linearly as clusters expand, effectively eliminating network bottlenecks and accelerating AI training.
“Our datasets span multiple regions, and we can’t afford to have our training schedule dictated by cross-region retrieval delays,” said Cécile Robert-Michon, director of internal infrastructure at Cohere. “CoreWeave AI Object Storage gives us a unified dataset footprint across regions with reads cached locally, so nothing waits on the network. It’s the difference between planning around our data and simply training.”
New capabilities now allow customers to run workloads across two regions while avoiding cross-region write latency and reducing long-term storage costs. Checkpoints can now be written locally within a region, closest to available compute. Data is written at local latency while being migrated to a second remote region in the background. This benefits organizations with compute distributed across multiple regions but data centralized in one, particularly those writing checkpoints during training jobs. The cross-region write feature minimizes downtime and enables the CKS cluster to continue training uninterrupted. Because the application interacts with a single bucket, no code changes are required, and permissions and retention rules function identically regardless of the originating region. This streamlines operations by removing the need to manually copy or relocate data for critical AI workloads.
Retaining critical data longer has also become more cost-effective. Teams frequently delete valuable data—such as the checkpoint from a near-successful run, datasets required to reproduce results, or model versions that may be referenced months later. The new Archive tier addresses this need with lower-cost storage specifically designed for archival data. It carries no fees for retrieval, early deletion, or reading directly from the Archive tier.
CoreWeave consistently delivers industry-leading performance, demonstrated by record-breaking MLPerf benchmark results in both inference and training. The company holds the distinction of being the only AI cloud to earn the top Platinum ranking in both SemiAnalysis ClusterMAX 1.0 and 2.0, and secured the #1 ranking for inference speed and price-performance for Moonshot AI’s Kimi K2.6 and Kimi K2.7 Code in independent inference benchmarking conducted by Artificial Analysis.
About CoreWeave
CoreWeave is The Essential Cloud for AI. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to move at the pace of innovation, building and scaling AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave serves as a force multiplier by combining superior infrastructure performance with deep technical expertise to accelerate breakthroughs. Established in 2017, CoreWeave completed its public listing on Nasdaq (CRWV) in March 2025. Learn more at www.coreweave.com.