# Microsoft Takes Delivery of the First Production Vera Rubin Racks as Nvidia Declares Full Production

The NVL72 configuration pairs Vera central processors with Rubin accelerators carrying up to 288 gigabytes of high-bandwidth memory, and Nvidia claims up to five times the inference throughput of Blackwell.

- Published: 2026-08-24T07:20:23.472Z
- Canonical: https://polylog.news/ai/2026-08-24/microsoft-takes-delivery-of-the-first-production-vera-rubin
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [NVIDIA Newsroom](https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory), [Microsoft Azure Blog](https://azure.microsoft.com/en-us/blog/microsofts-strategic-ai-datacenter-planning-enables-seamless-large-scale-nvidia-rubin-deployments/)

Satya Nadella, Microsoft's chief executive, published photographs from Microsoft data centers showing the first production Nvidia Vera Rubin racks in place. Nvidia then confirmed that the platform has [entered full-scale production](https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory). The delivery was relayed through [technical channels on 23 August](https://t.me/ai_machinelearning_big_data/10748) and dated to 21 August.

Vera Rubin is the successor to Blackwell. The NVL72 rack couples Vera central processing units with Rubin graphics processors that carry up to 288 gigabytes of fourth-generation high-bandwidth memory (HBM4). Nvidia states the configuration reaches up to five times the inference throughput of the previous generation. Racks are also running at CoreWeave, Google Cloud, Oracle Cloud Infrastructure and Nebius, so this is a broad rollout rather than a single-customer preview.

Microsoft has described its next-generation Fairwater sites as designed to [scale to hundreds of thousands of Vera Rubin superchips](https://azure.microsoft.com/en-us/blog/microsofts-strategic-ai-datacenter-planning-enables-seamless-large-scale-nvidia-rubin-deployments/), with the site design finished before the silicon arrived. That sequencing is the part engineers should note. The constraint on deploying a new rack generation is no longer fabrication yield alone. It is power delivery, liquid cooling loops and interconnect topology that had to be committed to concrete eighteen months earlier.

The performance-per-watt claim is Nvidia's own and has not been independently reproduced on public serving workloads. What is verified is the schedule: silicon that was announced as a roadmap item is now in racks at five named operators, which is a faster deployment pace than the Hopper-to-Blackwell transition.

## What this means

A five-fold inference throughput claim, even if it turns out to be only half that in practice, changes the unit economics of serving reasoning models. The parties exposed are the operators who signed multi-year capacity contracts on Blackwell-class hardware at prices set when tokens were scarcer. Cloud providers with the newest fleet can price serving below competitors running prior-generation racks, which pushes the inference price floor down again. The binding constraint moves further toward electricity and cooling, because a denser rack concentrates more load per square metre, and utilities in the regions hosting Fairwater-class sites are the ones who must find that supply.

## What to watch

- Independent throughput and cost-per-million-token measurements on Vera Rubin instances from serving vendors, which would show whether the five-times figure holds outside Nvidia's own testing.
- Whether cloud providers cut published inference prices in the quarter after Vera Rubin capacity comes online, the practical test of whether hardware gains reach customers or stay as margin.
- Power interconnection approvals and generation contracts around the new Microsoft sites, since a rack shipment that cannot be energised is not capacity.
