# Nvidia Puts Groq 3 LPX Into Full Production and Claims 30x More Agent Throughput Per Megawatt

The company says each liquid-cooled rack holds 256 language processing units, with the cloud provider Nebius as the first customer and racks online before the end of 2026.

- Published: 2026-08-25T06:26:21.272Z
- Canonical: https://polylog.news/ai/2026-08-25/nvidia-puts-groq-3-lpx-into-full-production-and-claims-30x-m
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [NVIDIA Blog (Groq 3 LPX, NVLink Fusion, Spectrum-X)](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/), [NVIDIA Blog (Vera Rubin NVL72 efficiency)](https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/), [CNBC](https://www.cnbc.com/2026/08/24/nvidia-says-groq-racks-will-be-online-this-year-after-20-billion-deal.html)

Nvidia used the Hot Chips 2026 conference to move its inference roadmap from announcement to shipment. The company said its [Groq 3 LPX accelerator has entered full production](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/), a latency-focused extension of the Vera Rubin data center platform, eight months after it signed a $20 billion licensing agreement with the chip designer Groq. [CNBC reported](https://www.cnbc.com/2026/08/24/nvidia-says-groq-racks-will-be-online-this-year-after-20-billion-deal.html) that the first racks will be online before the end of this year, deployed alongside Vera central processing units and Rubin graphics processing units (GPUs) at Nebius, a neocloud (a cloud-computing provider that specializes in renting out AI infrastructure), through its Token Factory platform.

The architecture divides the inference task into two parts. Vera Rubin NVL72 handles large-scale context processing. The Groq 3 LPX rack, which [StorageReview describes](https://www.storagereview.com/news/nvidia-groq-3-lpx-enters-full-production-3400-tokens-per-second-at-100k-context-256-lp30s-per-rack) as 256 language processing units (LPUs) per liquid-cooled rack reaching 3,400 tokens per second at a 100,000-token context, handles the latency-critical work of decoding tokens. Nvidia told CNBC that an LPX rack delivers up to 35 times more inference throughput per megawatt. The system also includes Spectrum-X Multiplane, an Ethernet fabric Nvidia says scales to 512,000 GPUs, BlueField-4 data processing units, and NVLink Fusion, a technology that lets third-party custom accelerators attach to Nvidia's sixth-generation NVLink rack systems.

The headline efficiency figure requires scrutiny. Nvidia says Vera Rubin NVL72 delivers [up to 30 times higher throughput per megawatt](https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/) than GB300 NVL72 on agentic workloads, but that figure comes from Nvidia running the SemiAnalysis AgentX workload itself, not from an outside lab. Nvidia attributes the gain to system-level engineering rather than the chip alone: mixture-of-experts serving runtimes (SGLang, TensorRT-LLM, vLLM), DeepGEMM kernels, the MXFP4 and MXFP8 mixed-precision formats, and the Dynamo session-aware serving stack.

The motivation is the shape of demand. Nvidia cites OpenRouter data showing that agentic workloads consume 15 times more tokens than a simple chat request, because each step's accumulated output becomes the input for the next step. That makes long-context prefill and fast decoding the two costs that determine whether an agent product is profitable.

## What this means

Nvidia has turned the Groq acquisition into shipped hardware within one year, and it is now selling throughput per megawatt rather than peak computing speed. That shift favors operators whose main constraint is electricity rather than capital, which now describes most large data center projects. The companies most exposed are AMD and the custom-chip programs at hyperscale cloud providers, which must now compete on tokens delivered per watt across a full rack system, not on chip specifications alone. Neoclouds such as Nebius gain a differentiated service they can price against commodity GPU rental, and companies building AI agent products gain a path to lower cost per unit of output that does not depend on a new model.

## What to watch

- Whether any party outside Nvidia reproduces the 30-times-per-megawatt claim on AgentX, since a vendor-run measurement of a vendor-selected workload is the weakest form of evidence for a number this large.
- How quickly Nebius and other neocloud providers publish real serving prices for LPX capacity, which would show whether the power efficiency actually reaches customers as a lower cost per token.
- Whether NVLink Fusion attracts custom accelerators from hyperscalers, which would signal Nvidia is willing to host rivals inside its rack in exchange for keeping control of the fabric.
