Morning Edition · Thursday, August 27, 2026Published at 2:20 AM EDT · New York
The 700-watt part claims 1.5 to 1.9 times the throughput per kilowatt of Nvidia's GB300, on a suite the company chose and ran itself.

OpenAI released the first performance results for Jalapeño, the inference-only accelerator it developed with Broadcom. Across three models, the company reports 1.5 to 1.9 times higher peak mixed tokens per second per kilowatt and 1.7 to 3.6 times lower end-to-end latency than the Nvidia GB200 and GB300 systems it selected for comparison. On Kimi K2.5, the largest model in the set, it claims roughly 1.5 times the performance per watt and 3.4 times lower latency.
The power figures are the substance of the claim. Jalapeño is a 700-watt package against 1,400 watts for the GB300, and Tom's Hardware reports that measured draw stayed at or below 550 watts during testing. In a buildout where electricity and rack density set the upper limit on deployable capacity, roughly half the package power for comparable throughput changes how much serving capacity fits behind a given substation.
The caveats are real and OpenAI states most of them. Jalapeño cannot train. The comparison uses SemiAnalysis's public InferenceX suite but the runs were executed and reported by OpenAI rather than by an independent lab. The target is also Nvidia's Blackwell-generation systems, not Vera Rubin, which has only just begun shipping. An application-specific integrated circuit built for one workload beating a general-purpose graphics processing unit on that workload is the expected result, not a surprise.
What matters commercially is not whether Jalapeño beats Nvidia in the abstract. It is that OpenAI, the largest single buyer of inference capacity, now has a credible in-house alternative for serving its own traffic and a published number to cite in every future procurement conversation.
OpenAI and Broadcom gain a published figure to cite in every future Nvidia procurement negotiation, and the claim lands while OpenAI is raising capital against a buildout whose economics depend on cheaper inference.
Part of a tracked trend
Custom Inference Silicon Displaces Merchant GPUs
The largest AI operators keep moving steady-state inference onto in-house accelerators designed with merchant silicon partners, splitting the accelerator market into a training near-monopoly and a contested, price-competitive inference tier.
Start a discussion in Townsquare.
More from this edition
The suite is SemiAnalysis's public InferenceX and the results were presented at Hot Chips, but OpenAI selected the models, ran the tests and chose Blackwell-generation systems rather than Vera Rubin as the comparison, and the chip has not yet been deployed at scale.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Nvidia's exposure here is on price rather than volume. The largest AI buyers, OpenAI with Broadcom, Google with its tensor processing units and Amazon with Trainium, are all moving steady-state inference onto internal silicon while continuing to buy merchant graphics processors for training and for burst capacity. That splits the market: training stays a near-monopoly, inference becomes contested, and inference is where token volume and therefore long-run spend concentrate. Broadcom is the direct beneficiary, since it captures the custom application-specific integrated circuit design and packaging revenue that would otherwise go to Nvidia's margin.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · OpenAI · Tom's Hardware
Comments
0No comments yet.