# OpenAI Publishes First Jalapeño Benchmarks, Claiming Up to 1.9 Times Nvidia's Throughput Per Kilowatt

The 700-watt inference chip, co-developed with Broadcom, was measured against Nvidia rack systems rated at 1,200 and 1,400 watts on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

- Published: 2026-08-26T06:22:56.640Z
- Canonical: https://polylog.news/ai/2026-08-26/openai-publishes-first-jalape-o-benchmarks-claiming-up-to-1
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [OpenAI](https://openai.com/index/jalapeno-first-results), [Polylog editors](https://polylog.news), [Tom's Hardware](https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks), [TechCrunch](https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/)

OpenAI released the [first published performance results](https://openai.com/index/jalapeno-first-results) for Jalapeño, the inference application-specific integrated circuit (ASIC) it co-developed with Broadcom. Across three models, GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, the company reports 1.5 to 1.9 times more throughput per kilowatt at peak and 1.7 to 3.6 times lower end-to-end latency than the best commercially available systems it tested. For interactive serving, where a single user waits on a response rather than a batch job, OpenAI puts the advantage at 2.1 to 4.1 times.

The power comparison is the part engineers should read closely. [Tom's Hardware reported](https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks) that the 700-watt Jalapeño part was measured against Nvidia accelerators rated at 1,200 and 1,400 watts, so a per-kilowatt win partly reflects each chip's power rating rather than a claim that Jalapeño beats the competing hardware on raw output per package. A chip built for one company's serving stack, with one tokenizer, one batching policy and one set of quantization choices, has structural advantages a general-purpose GPU cannot exploit.

The numbers are vendor-published, which is the standard caveat. What sets them apart from a typical lab blog post is the testing method: the runs used SemiAnalysis's public InferenceX benchmark, and SemiAnalysis verified some of them on site. That is partial third-party attestation, not independent reproduction. No outside party has yet run its own workloads on Jalapeño silicon, and OpenAI is not selling the chip. It plans to [deploy it in its own data centers](https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/) later this year.

The Russian-language technical channel AI ML Big Data [circulated the same figures](https://t.me/ai_machinelearning_big_data/10775) within hours, an indication of how closely non-US engineering audiences track the cost side of frontier serving. Market reaction was contained. [Seeking Alpha reported](https://seekingalpha.com/news/4636641-openais-jalapeno-spices-up-chip-market-as-it-outperforms-nvidias-blackwell) that investors did not appear alarmed, with Nvidia down just under 2 percent and Broadcom higher on the session.

Jalapeño was [unveiled with Broadcom in June](https://www.cnbc.com/2026/06/24/openai-and-broadcom-reveal-jalapeno-first-ai-chip-in-partnership.html). The pattern it belongs to is now well established: Google with tensor processing units, Amazon with Trainium, Meta with its own accelerators, and now OpenAI. Each removes a share of inference demand from the merchant market permanently, because the chip exists to serve one buyer's traffic and cannot be resold to anyone else.

## What this means

Inference, not training, is where the volume is, and it is the part of the workload most amenable to a narrow ASIC. If Jalapeño's per-watt advantage holds at rack scale, OpenAI lowers the marginal cost of serving reasoning models and gains room to cut API prices without cutting margin, which pressures every vendor reselling Nvidia capacity at a markup. Nvidia's exposure is not to a single competitor but to the largest AI buyers converting from customers into their own suppliers for the highest-volume workload. Broadcom gains through the opposite channel: it captures the design and packaging revenue that a merchant GPU sale would have carried.

## What to watch

- Whether any party outside OpenAI and SemiAnalysis publishes Jalapeño measurements on its own workloads. Independent reproduction is what separates a real efficiency advance from a favorable test configuration.
- How much of OpenAI's own serving traffic moves onto Jalapeño once deployment starts, and whether OpenAI keeps buying Nvidia at the same pace while it does. A flat Nvidia order book alongside rising internal silicon would show the chip is additive capacity, not a substitution.
- Whether OpenAI's published API prices fall in the months after deployment. A price cut would be the clearest evidence the cost advantage is real rather than a benchmark artifact.
