CoreWeave Puts Seven NVIDIA Vera Rubin NVL72 Racks Into Production

Dedicated Servers and GPU

CoreWeave has put seven NVIDIA Vera Rubin NVL72 racks into production across two regions, creating a 504-GPU deployment as the AI cloud operator moves NVIDIA’s latest rack-scale platform beyond single-rack operation.

The deployment, disclosed on September 16, connects multiple NVL72 systems into scale-out infrastructure designed to run training and inference workloads across hundreds of Rubin GPUs. CoreWeave said the systems are now operating on CoreWeave Cloud, alongside new storage capabilities intended to support increasingly distributed AI workloads.

Production deployment

Seven Vera Rubin NVL72 racks across two regions represent 504 NVIDIA Rubin GPUs. The infrastructure extends CoreWeave’s earlier single-rack Vera Rubin deployment into a multi-rack production environment.

From one NVL72 rack to a multi-rack system

CoreWeave brought up and validated its first NVIDIA Vera Rubin NVL72 rack in June. The latest deployment changes the infrastructure problem by requiring compute, networking, cooling, power and storage to operate consistently across several rack-scale systems rather than inside one NVLink domain.

Each Vera Rubin NVL72 rack integrates 72 Rubin GPUs and 36 Vera CPUs. The rack also incorporates NVIDIA NVLink 6, ConnectX-9 SuperNICs, BlueField-4 DPUs, high-speed storage and liquid cooling.

Deployment detail Confirmed configuration
Production racks 7 NVIDIA Vera Rubin NVL72 racks
Total Rubin GPUs 504
GPUs per rack 72
CPUs per rack 36 NVIDIA Vera CPUs
Deployment footprint Two regions
Backend connectivity Up to 1.6 Tbps per GPU through two ConnectX-9 SuperNICs

Networking becomes part of the compute system

Inside an NVL72 rack, NVLink 6 provides the high-bandwidth scale-up fabric. Once workloads extend across racks, the backend network has to carry synchronized traffic between separate NVLink domains without leaving expensive accelerators waiting for data.

CoreWeave says its Vera Rubin architecture uses a two-tier, non-blocking, multi-rail and multi-plane RoCE network. Each Rubin GPU is served by two ConnectX-9 SuperNICs, providing up to 1.6 Tbps of backend network connectivity per GPU.

The company also tests the scale-out fabric independently from the fastest GPU-to-GPU paths. During validation, CoreWeave can disable those faster links and force traffic over the backend network to identify problems involving switches, cabling, congestion, latency or configuration before systems enter production.

This is an important distinction for multi-rack AI infrastructure. Adding accelerators does not by itself guarantee that a distributed training or inference job will scale efficiently. Network bottlenecks, a slow GPU or an unreliable connection can hold back a workload spanning otherwise healthy racks.

Cooling and power move into the operational layer

The same principle applies to physical infrastructure. CoreWeave said it extended its Mission Control software with rack-management and programmable liquid-cooling systems that allow compute, power delivery, cooling and environmental conditions to be managed as parts of the same rack lifecycle.

The Vera Rubin systems use liquid cooling, and CoreWeave’s technical description specifies a 45°C cooling loop. Its Valvey system provides programmable control of cooling flow at the rack level, while its Racky management layer and Rack LifeCycle Controller coordinate rack operations.

That integration matters because a thermal or power issue can affect a distributed workload even when the GPUs themselves remain functional. At multi-rack scale, infrastructure health increasingly has to be evaluated at the system level rather than component by component.

Storage is being scaled with the GPUs

CoreWeave also announced two changes to its AI Object Storage platform alongside the Vera Rubin deployment: cross-region write acceleration and a new Archive tier.

The company is using a Local Object Transport Accelerator, or LOTA, to provide a compute-local NVMe endpoint for cached object data. CoreWeave says the system can deliver up to 7 GBps per GPU under appropriate workload and cache conditions.

Cross-region write acceleration is designed for operations such as AI checkpointing. Data can be written durably and acknowledged in the local region while replication to another region continues in the background. CoreWeave says that migration typically completes within 72 hours.

Why the multi-rack milestone matters

The September deployment provides a more meaningful infrastructure test than operating a single NVL72 rack. A single rack validates NVIDIA’s tightly integrated rack-scale architecture; a multi-rack deployment tests whether a cloud operator can make several of those systems behave as a larger production cluster.

For operators deploying the next generation of AI infrastructure, that shifts engineering attention toward the systems surrounding the GPUs: scale-out networking, liquid cooling, power management, storage throughput, monitoring and automated validation.

CoreWeave says its current network architecture is designed to extend to approximately 128,000 GPUs. That figure describes the architecture’s intended scaling capability, not the size of the Vera Rubin deployment announced this week. The confirmed production footprint disclosed by independent reporting is seven racks and 504 Rubin GPUs.

What happens next

The immediate question is how quickly multi-rack Vera Rubin deployments expand beyond the initial seven systems and how the architecture performs under sustained customer workloads.

CoreWeave has not disclosed a 128,000-GPU Rubin deployment. For now, the concrete change is smaller but significant for infrastructure operators: Vera Rubin has moved at CoreWeave from a validated single rack to production infrastructure spanning multiple NVL72 systems and hundreds of GPUs.