# OpenAI Publishes First Benchmarks for Its Jalapeño Inference Chip Against Nvidia Systems

The 700-watt part claims 1.5 to 1.9 times the throughput per kilowatt of Nvidia's GB300, on a suite the company chose and ran itself.

- Published: 2026-08-27T06:20:08.967Z
- Canonical: https://polylog.news/ai/2026-08-27/openai-publishes-first-benchmarks-for-its-jalape-o-inference
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Polylog editors](https://polylog.news), [OpenAI](https://openai.com/index/jalapeno-first-results/), [Tom's Hardware](https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks)

OpenAI released the [first performance results for Jalapeño](https://openai.com/index/jalapeno-first-results/), the inference-only accelerator it developed with Broadcom. Across three models, the company reports 1.5 to 1.9 times higher peak mixed tokens per second per kilowatt and 1.7 to 3.6 times lower end-to-end latency than the Nvidia GB200 and GB300 systems it selected for comparison. On Kimi K2.5, the largest model in the set, it claims roughly 1.5 times the performance per watt and 3.4 times lower latency.

The power figures are the substance of the claim. Jalapeño is a 700-watt package against 1,400 watts for the GB300, and [Tom's Hardware reports](https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks) that measured draw stayed at or below 550 watts during testing. In a buildout where electricity and rack density set the upper limit on deployable capacity, roughly half the package power for comparable throughput changes how much serving capacity fits behind a given substation.

The caveats are real and OpenAI states most of them. Jalapeño cannot train. The comparison uses SemiAnalysis's public InferenceX suite but the runs were executed and reported by OpenAI rather than by an independent lab. The target is also Nvidia's Blackwell-generation systems, not Vera Rubin, which has only just begun shipping. An application-specific integrated circuit built for one workload beating a general-purpose graphics processing unit on that workload is the expected result, not a surprise.

What matters commercially is not whether Jalapeño beats Nvidia in the abstract. It is that OpenAI, the largest single buyer of inference capacity, now has a credible in-house alternative for serving its own traffic and a published number to cite in every future procurement conversation.

## What this means

Nvidia's exposure here is on price rather than volume. The largest AI buyers, OpenAI with Broadcom, Google with its tensor processing units and Amazon with Trainium, are all moving steady-state inference onto internal silicon while continuing to buy merchant graphics processors for training and for burst capacity. That splits the market: training stays a near-monopoly, inference becomes contested, and inference is where token volume and therefore long-run spend concentrate. Broadcom is the direct beneficiary, since it captures the custom application-specific integrated circuit design and packaging revenue that would otherwise go to Nvidia's margin.

## What to watch

- Whether an independent party reruns the InferenceX comparison on the same hardware. Vendor-run benchmarks and third-party reproductions have diverged before, and confirmation would move this from a marketing claim to an engineering fact.
- How many Jalapeño units OpenAI actually deploys and what share of ChatGPT and application programming interface traffic moves onto them. Announced silicon and deployed silicon are different things, and the gap is where these projects usually fail.
- Whether Nvidia publishes Vera Rubin inference efficiency numbers in response. A rapid answer would show the company treats custom inference silicon as a genuine competitive threat rather than a niche.
