Morning Edition · Wednesday, August 26, 2026Published at 2:22 AM EDT · New York
The 700-watt inference chip, co-developed with Broadcom, was measured against Nvidia rack systems rated at 1,200 and 1,400 watts on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

OpenAI released the first published performance results for Jalapeño, the inference application-specific integrated circuit (ASIC) it co-developed with Broadcom. Across three models, GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, the company reports 1.5 to 1.9 times more throughput per kilowatt at peak and 1.7 to 3.6 times lower end-to-end latency than the best commercially available systems it tested. For interactive serving, where a single user waits on a response rather than a batch job, OpenAI puts the advantage at 2.1 to 4.1 times.
The power comparison is the part engineers should read closely. Tom's Hardware reported that the 700-watt Jalapeño part was measured against Nvidia accelerators rated at 1,200 and 1,400 watts, so a per-kilowatt win partly reflects each chip's power rating rather than a claim that Jalapeño beats the competing hardware on raw output per package. A chip built for one company's serving stack, with one tokenizer, one batching policy and one set of quantization choices, has structural advantages a general-purpose GPU cannot exploit.
The numbers are vendor-published, which is the standard caveat. What sets them apart from a typical lab blog post is the testing method: the runs used SemiAnalysis's public InferenceX benchmark, and SemiAnalysis verified some of them on site. That is partial third-party attestation, not independent reproduction. No outside party has yet run its own workloads on Jalapeño silicon, and OpenAI is not selling the chip. It plans to deploy it in its own data centers later this year.
The Russian-language technical channel AI ML Big Data circulated the same figures within hours, an indication of how closely non-US engineering audiences track the cost side of frontier serving. Market reaction was contained. Seeking Alpha reported that investors did not appear alarmed, with Nvidia down just under 2 percent and Broadcom higher on the session.
Jalapeño was unveiled with Broadcom in June. The pattern it belongs to is now well established: Google with tensor processing units, Amazon with Trainium, Meta with its own accelerators, and now OpenAI. Each removes a share of inference demand from the merchant market permanently, because the chip exists to serve one buyer's traffic and cannot be resold to anyone else.
Part of a tracked trend
Frontier Labs Build Their Own Inference Silicon
The largest AI operators keep converting from buyers of merchant accelerators into designers of their own inference chips, permanently removing their highest-volume workloads from the general-purpose GPU market and shifting value toward custom silicon design partners.
Start a discussion in Townsquare.
More from this edition
OpenAI and Broadcom, which gain leverage over Nvidia pricing and a case for custom silicon, while Nvidia's merchant-GPU margin story absorbs the doubt.
The numbers are real and partly witnessed by SemiAnalysis, but SemiAnalysis itself calls the Blackwell comparison "somewhat incomplete and unfair": Jalapeño ran single-token prediction against Nvidia configurations using multi-token prediction, and it was never tested against Vera Rubin, the platform OpenAI has itself agreed to deploy.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Inference, not training, is where the volume is, and it is the part of the workload most amenable to a narrow ASIC. If Jalapeño's per-watt advantage holds at rack scale, OpenAI lowers the marginal cost of serving reasoning models and gains room to cut API prices without cutting margin, which pressures every vendor reselling Nvidia capacity at a markup. Nvidia's exposure is not to a single competitor but to the largest AI buyers converting from customers into their own suppliers for the highest-volume workload. Broadcom gains through the opposite channel: it captures the design and packaging revenue that a merchant GPU sale would have carried.
What to watch
Observations to monitor, not financial advice.
Synthesized from: OpenAI · Polylog editors · Tom's Hardware · TechCrunch
Comments
1Aug 27, 3:38 AM · edited
Against the 1,400 watt baseline, a 1.9x throughput per kilowatt ratio implies Jalapeño delivers roughly 95 percent of that system's absolute throughput at half the power draw, so the gain is efficiency and latency, not aggregate output.