CoreWeave deploys multi-rack Nvidia Vera Rubin NVL72 cluster, stock jumps 14%

Photo: Matheus Bertelli / Pexels

CoreWeave deploys multi-rack Nvidia Vera Rubin NVL72 cluster, stock jumps 14%

The AI cloud provider says it's the first to operationalize Nvidia's next-generation GPU system, claiming 10x efficiency gains over Blackwell

CoreWeave just became the first AI cloud provider to get Nvidia’s Vera Rubin NVL72 system up and running, connecting hundreds of next-generation GPUs into a production cluster designed for the most demanding AI workloads on the planet.

The market noticed. CoreWeave’s stock (CRWV) surged roughly 14% on the news, with trading volume nearly double its three-month average.

What’s inside the rack

Dell delivered the NVL72 hardware on May 31, and CoreWeave had the system operational by June 1. The turnaround time between receiving a cutting-edge GPU rack and actually running workloads on it was, in other words, about 24 hours.

Each NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs, all connected through NVLink 6 interconnect running at 260 TB/s of bandwidth.

Advertisement

The system delivers up to 3.6 exaFLOPS of NVFP4 inference capability per rack. CoreWeave claims the Rubin system offers 10 times the inference throughput per megawatt compared to the previous Blackwell generation. It also reportedly halves the number of GPUs needed for equivalent workloads and drops the cost to roughly one-tenth per million tokens versus Blackwell.

The architecture supports scaling well beyond a single rack. CoreWeave says clusters can expand to thousands of GPUs using either Spectrum-X Ethernet or Quantum-X800 InfiniBand networking.

The software layer matters too

CoreWeave also announced proprietary software tools built to manage the physical complexity of running liquid-cooled, ultra-dense GPU racks at scale.

Valvey, its liquid cooling management system, provides real-time monitoring of thermal performance across racks. Racky handles unified rack control, giving operators a single interface to manage the complex interplay of power, cooling, and compute across an entire deployment.

On the storage side, CoreWeave introduced LOTA, a high-performance storage solution that achieves data transport speeds of up to 7 GB/s per GPU.

Why being first matters

CoreWeave has built its entire business model around being faster than the hyperscalers at deploying Nvidia’s latest silicon. Being the first cloud provider to operationalize Vera Rubin NVL72 extends a pattern. The company has consistently been among the earliest adopters of each new Nvidia generation, giving it a window of exclusivity that attracts AI companies willing to pay premium prices for access to the newest hardware.

CoreWeave expects full production availability of Rubin clusters later in 2026, with major hyperscalers also planning adoption.

What to watch from here

The 10x efficiency improvement and cost reduction figures come from CoreWeave’s own benchmarks. The 4x faster training claim for certain scenarios is worth monitoring closely. If validated by independent benchmarks, it could accelerate the timeline for next-generation AI models and increase demand for Rubin clusters beyond what the current supply chain can handle.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
CoreWeave deploys multi-rack Nvidia Vera Rubin NVL72 cluster, stock jumps 14%
CoreWeave deploys multi-rack Nvidia Vera Rubin NVL72 cluster, stock jumps 14%

The AI cloud provider says it's the first to operationalize Nvidia's next-generation GPU system, claiming 10x efficiency gains over Blackwell

Photo: Matheus Bertelli / Pexels

CoreWeave just became the first AI cloud provider to get Nvidia’s Vera Rubin NVL72 system up and running, connecting hundreds of next-generation GPUs into a production cluster designed for the most demanding AI workloads on the planet.

The market noticed. CoreWeave’s stock (CRWV) surged roughly 14% on the news, with trading volume nearly double its three-month average.

What’s inside the rack

Dell delivered the NVL72 hardware on May 31, and CoreWeave had the system operational by June 1. The turnaround time between receiving a cutting-edge GPU rack and actually running workloads on it was, in other words, about 24 hours.

Each NVL72 rack packs 72 Rubin GPUs and 36 Vera CPUs, all connected through NVLink 6 interconnect running at 260 TB/s of bandwidth.

Advertisement

The system delivers up to 3.6 exaFLOPS of NVFP4 inference capability per rack. CoreWeave claims the Rubin system offers 10 times the inference throughput per megawatt compared to the previous Blackwell generation. It also reportedly halves the number of GPUs needed for equivalent workloads and drops the cost to roughly one-tenth per million tokens versus Blackwell.

The architecture supports scaling well beyond a single rack. CoreWeave says clusters can expand to thousands of GPUs using either Spectrum-X Ethernet or Quantum-X800 InfiniBand networking.

The software layer matters too

CoreWeave also announced proprietary software tools built to manage the physical complexity of running liquid-cooled, ultra-dense GPU racks at scale.

Valvey, its liquid cooling management system, provides real-time monitoring of thermal performance across racks. Racky handles unified rack control, giving operators a single interface to manage the complex interplay of power, cooling, and compute across an entire deployment.

On the storage side, CoreWeave introduced LOTA, a high-performance storage solution that achieves data transport speeds of up to 7 GB/s per GPU.

Why being first matters

CoreWeave has built its entire business model around being faster than the hyperscalers at deploying Nvidia’s latest silicon. Being the first cloud provider to operationalize Vera Rubin NVL72 extends a pattern. The company has consistently been among the earliest adopters of each new Nvidia generation, giving it a window of exclusivity that attracts AI companies willing to pay premium prices for access to the newest hardware.

CoreWeave expects full production availability of Rubin clusters later in 2026, with major hyperscalers also planning adoption.

What to watch from here

The 10x efficiency improvement and cost reduction figures come from CoreWeave’s own benchmarks. The 4x faster training claim for certain scenarios is worth monitoring closely. If validated by independent benchmarks, it could accelerate the timeline for next-generation AI models and increase demand for Rubin clusters beyond what the current supply chain can handle.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.