# Custom Inference Silicon Displaces Merchant GPUs

The largest AI operators keep moving steady-state inference onto in-house accelerators designed with merchant silicon partners, splitting the accelerator market into a training near-monopoly and a contested, price-competitive inference tier.

- Conviction: 40 / 100 (forming)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-08-27T00:00:00.000Z
- Last updated: 2026-08-27T14:00:31.454Z
- Canonical: https://polylog.news/ai/trends/custom-inference-silicon-displaces-merchant-gpus
- Publisher: Polylog
- Affected regions: United States, China

## Recent score history

- 2026-08-27: 40
- 2026-08-28: 34

## Recent evidence

- [confirms] Nvidia Opens Its Custom Memory Design to Rival Accelerators Through NVLink Fusion (2026-08-27): Nvidia is opening NVHBM — a custom memory design moving the controller into the HBM stack for a claimed 30 percent more bandwidth and 15 percent lower memory power than HBM4E — to rival accelerators via NVLink Fusion, with Amazon's Annapurna Labs as the first partner. Nvidia selling its memory advantage to the in-house silicon team most likely to displace its GPUs is an incumbent conceding that the inference tier will be contested and choosing to be paid inside it.
- [confirms] OpenAI Publishes First Benchmarks for Its Jalapeño Inference Chip Against Nvidia Systems (2026-08-27): OpenAI's Jalapeño claims 1.5–1.9x throughput per kilowatt versus Nvidia's GB300 on a suite OpenAI chose and ran itself. If the perf-per-watt gap holds under independent testing, the highest-volume steady-state inference workload at one of the largest buyers leaves the merchant GPU tier.
