Morning Edition · Tuesday, August 25, 2026Published at 2:26 AM EDT · New York
The gap comes mostly from serving software, and the analysts say Nvidia would still be cheaper per token even if the competing hardware were free.

SemiAnalysis published AgentX, an open-source benchmark that replays recorded coding-agent sessions instead of synthetic prompts, and the result is unfavorable for AMD. At a target interactivity of 150 output tokens per second per user, the analysts report that Nvidia hardware reaches up to five times better cost efficiency than AMD when serving the GLM 5.3 model through the open-source SGLang runtime. Forbes summarized the same finding, noting cost advantages as wide as 20 times on certain open-source software stacks.
The most significant finding in the analysis is not the multiple itself. SemiAnalysis argues that at 150 tokens per second per user, the gap is large enough that cost per token would still favor Nvidia even if a competitor gave its accelerators away for free, once hosting and power costs are counted. That is a claim about software quality, not about the chips themselves. The analysts attribute Nvidia's lead to cache management, request routing, and incremental tokenization, factors that matter far more under agent traffic than under single-turn chat.
The shape of the workload explains why. AgentX uses anonymized Claude Code sessions with a median input of roughly 142,000 tokens against a median output of 444 tokens, plus long idle gaps between turns. Under that profile, any stack that recomputes context aggressively or schedules requests poorly performs worse, while stacks that reuse prefixes and chunks perform better.
Two caveats limit how broadly this number applies. The comparison sets Nvidia's current Blackwell generation against AMD's current CDNA 4 chips, while Nvidia's upcoming Rubin generation, Google's tensor processing units, and AMD's upcoming MI455X chip are not yet included. AgentX also measures one category of work, coding agents, which happens to be the workload Nvidia's serving software has been tuned hardest against. A second AI Post summary of the report repeated the five-times figure without those qualifications, which shows how a benchmark result can turn into an unqualified market narrative.
Part of a tracked trend
Serving Software Decides Accelerator Economics
The cost gap between AI accelerators is increasingly set by scheduling, caching, and runtime quality rather than by silicon specifications, so software maturity becomes the durable moat and challengers keep losing on total cost even when their hardware specs are competitive.
Start a discussion in Townsquare.
More from this edition
Nvidia, whose pricing power in inference rests on the argument that a competitor's cheaper silicon cannot overcome its software deficit, and every investor holding the assumption that the CUDA ecosystem is a durable moat rather than a temporary lead.
The five-times figure and the "even if the hardware were free" line are stated directly in SemiAnalysis's AgentX report and reported accurately here, but it measures a software snapshot at one interactivity target on one workload class, Nvidia itself publishes an operator guide for running AgentX, and no MI455X or Rubin parts are in the comparison, so the result describes today's ROCm software state rather than a fixed hardware gap.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
The competitive question for AI accelerators has shifted from peak throughput to how well the serving software handles long, repetitive, bursty agent workloads. That favors Nvidia's CUDA ecosystem, where kernel and runtime engineering has accumulated advantages for years, and it puts pressure on AMD to close a software gap that new hardware alone will not close. Enterprises evaluating alternative accelerators now have a concrete public methodology to test claims themselves, which requires stronger evidence from vendors on both sides. If AMD's ROCm software improves measurably on AgentX before the MI455X ships, pricing pressure on Nvidia's inference capacity will return. If it does not, Nvidia keeps its pricing power in the fastest-growing segment of inference demand.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · NVIDIA Blog (NVLink Fusion and XPUs)
Comments
0No comments yet.