Morning Edition · Tuesday, August 25, 2026Published at 2:26 AM EDT · New York
The company says each liquid-cooled rack holds 256 language processing units, with the cloud provider Nebius as the first customer and racks online before the end of 2026.

Nvidia used the Hot Chips 2026 conference to move its inference roadmap from announcement to shipment. The company said its Groq 3 LPX accelerator has entered full production, a latency-focused extension of the Vera Rubin data center platform, eight months after it signed a $20 billion licensing agreement with the chip designer Groq. CNBC reported that the first racks will be online before the end of this year, deployed alongside Vera central processing units and Rubin graphics processing units (GPUs) at Nebius, a neocloud (a cloud-computing provider that specializes in renting out AI infrastructure), through its Token Factory platform.
The architecture divides the inference task into two parts. Vera Rubin NVL72 handles large-scale context processing. The Groq 3 LPX rack, which StorageReview describes as 256 language processing units (LPUs) per liquid-cooled rack reaching 3,400 tokens per second at a 100,000-token context, handles the latency-critical work of decoding tokens. Nvidia told CNBC that an LPX rack delivers up to 35 times more inference throughput per megawatt. The system also includes Spectrum-X Multiplane, an Ethernet fabric Nvidia says scales to 512,000 GPUs, BlueField-4 data processing units, and NVLink Fusion, a technology that lets third-party custom accelerators attach to Nvidia's sixth-generation NVLink rack systems.
The headline efficiency figure requires scrutiny. Nvidia says Vera Rubin NVL72 delivers up to 30 times higher throughput per megawatt than GB300 NVL72 on agentic workloads, but that figure comes from Nvidia running the SemiAnalysis AgentX workload itself, not from an outside lab. Nvidia attributes the gain to system-level engineering rather than the chip alone: mixture-of-experts serving runtimes (SGLang, TensorRT-LLM, vLLM), DeepGEMM kernels, the MXFP4 and MXFP8 mixed-precision formats, and the Dynamo session-aware serving stack.
The motivation is the shape of demand. Nvidia cites OpenRouter data showing that agentic workloads consume 15 times more tokens than a simple chat request, because each step's accumulated output becomes the input for the next step. That makes long-context prefill and fast decoding the two costs that determine whether an agent product is profitable.
Part of a tracked trend
The Inference-Cost Efficiency Race
Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.
Start a discussion in Townsquare.
More from this edition
Nvidia, which converts a $20 billion Groq licensing deal into a shipping product and reframes competition around throughput per megawatt, a metric it currently defines and measures, along with Nebius, whose share price depends on being first to sell differentiated inference capacity.
Full production, the Nebius deployment and the 2026 timeline are confirmed by Nvidia's own newsroom and CNBC, but the efficiency number is not: Nvidia states the 30-times figure is its own AgentX measurement pending SemiAnalysis review and excludes Vera central processor tool-calling performance, and the separate 35-times figure in circulation refers to token cost reduction rather than throughput per megawatt, so the article's "35 times more inference throughput per megawatt" merges two different claims.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Nvidia has turned the Groq acquisition into shipped hardware within one year, and it is now selling throughput per megawatt rather than peak computing speed. That shift favors operators whose main constraint is electricity rather than capital, which now describes most large data center projects. The companies most exposed are AMD and the custom-chip programs at hyperscale cloud providers, which must now compete on tokens delivered per watt across a full rack system, not on chip specifications alone. Neoclouds such as Nebius gain a differentiated service they can price against commodity GPU rental, and companies building AI agent products gain a path to lower cost per unit of output that does not depend on a new model.
What to watch
Observations to monitor, not financial advice.
Synthesized from: NVIDIA Blog (Groq 3 LPX, NVLink Fusion, Spectrum-X) · NVIDIA Blog (Vera Rubin NVL72 efficiency) · CNBC
Comments
1Aug 26, 2:02 AM · edited
If the 30x per megawatt claim holds, energy cost per inference token falls roughly 97 percent versus the baseline, materially changing the unit economics of continuous agent deployments.