CoreWeave는 다중 랙 Nvidia Vera Rubin NVL72 클러스터를 성공적으로 가동했으며, Coreweave AI Object Storage 플랫폼에서 새로운 기능을 출시했다.
CoreWeave, Inc. announced the operational deployment of multi-rack NVIDIA Vera Rubin NVL72 clusters on CoreWeave Cloud, integrating hundreds of NVIDIA Rubin GPUs into a single scale-out architecture designed for agentic AI. The company also unveiled two new capabilities for CoreWeave AI Object Storage—cross-region write acceleration and a new Archive tier—ensuring that the data these workloads depend on remains closely coupled with the GPUs. These multi-rack Vera Rubin NVL72 configurations enable training and inference jobs to span hundreds of Rubin GPUs simultaneously.
A single NVIDIA Vera Rubin NVL72 rack pairs 72 Rubin GPUs with 36 Vera CPUs, alongside NVIDIA NVLink 6, NVIDIA ConnectX-9 SuperNICs, and NVIDIA BlueField-4 DPUs. To support distributed workloads, CoreWeave introduced cross-region write acceleration, which eliminates latency delays when operating across multiple geographic regions. As a job writes data locally, CoreWeave asynchronously replicates it to a secondary region in the background. This ensures that an agent’s intermediate states, retrieved contexts, and outputs advance at the speed of its reasoning.
Multi-rack Vera Rubin NVL72 deployments unify hundreds of accelerators across multiple racks into a single scale-out cluster using NVIDIA Spectrum-X Ethernet networking. This architecture delivers the capacity to train larger models, handle more demanding inference workloads, and execute reinforcement learning at scale. Achieving this requires precise engineering across every layer—including compute, networking, storage, cooling, power, firmware, and software—to function as a single coordinated system. CoreWeave orchestrates this integration by automating rack lifecycle control.
Upon arrival, racks require physical connection and rigorous validation. CoreWeave Mission Control automates the setup process through the Rack LifeCycle Controller, which coordinates hardware detection, firmware updates, validation, power distribution, and thermal management. Within this framework, Racky handles rack-level control while Valvey executes cooling actions. Performance validation spans from individual components to full-system benchmarks. CoreWeave combines NVIDIA field diagnostics with comprehensive full-rack workload testing, meticulously comparing every metric against established baselines. Leveraging years of real-world operational experience, the company ensures readiness before any rack enters production.
Only racks that meet these stringent system-level thresholds proceed to production, guaranteeing optimal GPU performance. Network scaling is tightly aligned with compute growth: each Rubin GPU is equipped with two NVIDIA ConnectX-9 SuperNICs, delivering 1.6 Tb/s of scale-out connectivity per GPU across multiplane, multirail paths. This non-blocking fabric supports approximately 128,000 GPUs per rail. Furthermore, the modular topology allows additional racks to be integrated without requiring fabric redesigns during expansion.
CoreWeave AI Object Storage LOTA (Local Object Transport Accelerator) further optimizes data proximity for AI workloads by implementing managed caching on every CoreWeave Kubernetes Service node. LOTA delivers read speeds matching local NVMe drives, reducing latency by 8x compared to traditional storage clusters. It provides up to 7 GB/s of throughput per GPU and scales linearly as clusters expand, effectively eliminating network bottlenecks and accelerating AI training cycles.
The newly launched cross-region write capability enables customers to distribute workloads across two regions while bypassing cross-region write latency and lowering long-term storage expenses. Checkpoints are now written locally within the region hosting the available compute, allowing data to be ingested at local latency while being asynchronously migrated to a secondary remote region. This feature minimizes training pauses, enabling CKS clusters to maintain continuous operation. Because the application interacts with a single logical bucket, no code modifications are required, and permissions and retention policies remain consistent regardless of the originating region. This streamlines operations by removing the need to manually transfer or replicate data for critical AI workloads.
The Archive tier is a cost-optimized storage class designed specifically for infrequently accessed data, featuring zero fees for retrieval, early deletion, or standard reads. CoreWeave continues to deliver industry-leading performance, evidenced by record-breaking MLPerf benchmark results in both inference and training. It remains the only AI cloud provider to achieve the top Platinum ranking in both SemiAnalysis ClusterMAX 1.0 and 2.0, and secured the #1 ranking for inference speed and price-performance for Moonshot AI’s Kimi K2.6 and Kimi K2.7 Code in independent benchmarking conducted by Artificial Analysis.