Custom Inference Silicon Displaces Merchant GPUs
The largest AI operators keep moving steady-state inference onto in-house accelerators designed with merchant silicon partners, splitting the accelerator market into a training near-monopoly and a contested, price-competitive inference tier.
forming · confidence 40 · Emerging (watchlist) · tracking since August 27, 2026 · updated August 27, 2026
Score history
Daily conviction score, 0 to 100. Higher means the thesis is more strongly corroborated.
Now 40 · -6 since Aug 27 · ranged 34 to 40
Why the conviction moved
- Aug 28Weakened
Amazon, the operator furthest along on in-house accelerators with Trainium, committed to two million additional Nvidia GPUs for 2027 and 2028 and put Annapurna Labs to work as Nvidia's first NVHBM memory partner rather than purely on displacement silicon. The largest custom-silicon builder deepening both its purchase volume and its engineering entanglement with the merchant vendor weakens the claim that steady-state inference is migrating off merchant GPUs.
- Aug 27Strengthened +5
OpenAI's Jalapeño claims 1.5–1.9x throughput per kilowatt versus Nvidia's GB300 on a suite OpenAI chose and ran itself. If the perf-per-watt gap holds under independent testing, the highest-volume steady-state inference workload at one of the largest buyers leaves the merchant GPU tier.
- Aug 27Strengthened +6
Nvidia is opening NVHBM — a custom memory design moving the controller into the HBM stack for a claimed 30 percent more bandwidth and 15 percent lower memory power than HBM4E — to rival accelerators via NVLink Fusion, with Amazon's Annapurna Labs as the first partner. Nvidia selling its memory advantage to the in-house silicon team most likely to displace its GPUs is an incumbent conceding that the inference tier will be contested and choosing to be paid inside it.
Source trail
Supporting · August 27, 2026
Nvidia Opens Its Custom Memory Design to Rival Accelerators Through NVLink Fusion
Nvidia is opening NVHBM — a custom memory design moving the controller into the HBM stack for a claimed 30 percent more bandwidth and 15 percent lower memory power than HBM4E — to rival accelerators via NVLink Fusion, with Amazon's Annapurna Labs as the first partner. Nvidia selling its memory advantage to the in-house silicon team most likely to displace its GPUs is an incumbent conceding that the inference tier will be contested and choosing to be paid inside it.
NVIDIA BlogSupporting · August 27, 2026
OpenAI Publishes First Benchmarks for Its Jalapeño Inference Chip Against Nvidia Systems
OpenAI's Jalapeño claims 1.5–1.9x throughput per kilowatt versus Nvidia's GB300 on a suite OpenAI chose and ran itself. If the perf-per-watt gap holds under independent testing, the highest-volume steady-state inference workload at one of the largest buyers leaves the merchant GPU tier.
AI Post (Telegram)
Unlock full source trail, score history, and daily updates.
Unlock TrendsAffected regions & assets
Townsquare
Argue the thesis in Townsquare.