Morning Edition · Monday, August 24, 2026Published at 3:20 AM EDT · New York
The NVL72 configuration pairs Vera central processors with Rubin accelerators carrying up to 288 gigabytes of high-bandwidth memory, and Nvidia claims up to five times the inference throughput of Blackwell.

Satya Nadella, Microsoft's chief executive, published photographs from Microsoft data centers showing the first production Nvidia Vera Rubin racks in place. Nvidia then confirmed that the platform has entered full-scale production. The delivery was relayed through technical channels on 23 August and dated to 21 August.
Vera Rubin is the successor to Blackwell. The NVL72 rack couples Vera central processing units with Rubin graphics processors that carry up to 288 gigabytes of fourth-generation high-bandwidth memory (HBM4). Nvidia states the configuration reaches up to five times the inference throughput of the previous generation. Racks are also running at CoreWeave, Google Cloud, Oracle Cloud Infrastructure and Nebius, so this is a broad rollout rather than a single-customer preview.
Microsoft has described its next-generation Fairwater sites as designed to scale to hundreds of thousands of Vera Rubin superchips, with the site design finished before the silicon arrived. That sequencing is the part engineers should note. The constraint on deploying a new rack generation is no longer fabrication yield alone. It is power delivery, liquid cooling loops and interconnect topology that had to be committed to concrete eighteen months earlier.
The performance-per-watt claim is Nvidia's own and has not been independently reproduced on public serving workloads. What is verified is the schedule: silicon that was announced as a roadmap item is now in racks at five named operators, which is a faster deployment pace than the Hopper-to-Blackwell transition.
Nvidia and Microsoft both gain from a delivery milestone that shows the roadmap holding on schedule, and cloud operators running the newest fleet can undercut competitors still amortising Blackwell-class racks.
Part of a tracked trend
Power Becomes the Binding Constraint on AI Buildout
Electricity availability, not chip supply, increasingly determines how fast AI capacity comes online, pushing the largest operators to finance and build their own generation and turning power equipment, fuel supply and permitting into recurring chokepoints with direct pricing consequences.
Start a discussion in Townsquare.
More from this edition
The "first production systems" description competes with CoreWeave's separate claim to the industry-first bring-up and validation of the same rack, and the five-times inference figure is Nvidia's own rack-level number that no independent party has reproduced on public serving workloads.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
A five-fold inference throughput claim, even if it turns out to be only half that in practice, changes the unit economics of serving reasoning models. The parties exposed are the operators who signed multi-year capacity contracts on Blackwell-class hardware at prices set when tokens were scarcer. Cloud providers with the newest fleet can price serving below competitors running prior-generation racks, which pushes the inference price floor down again. The binding constraint moves further toward electricity and cooling, because a denser rack concentrates more load per square metre, and utilities in the regions hosting Fairwater-class sites are the ones who must find that supply.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · NVIDIA Newsroom · Microsoft Azure Blog
Comments
1Aug 25, 1:58 AM · edited
At 72 Rubin GPUs times 288 GB each, a single NVL72 rack carries roughly 20.7 TB of HBM4 in aggregate, sufficient to hold simultaneous full precision copies of several frontier class models without quantization.