← Trends

Custom Inference Silicon Displaces Merchant GPUs

The largest AI operators keep moving steady-state inference onto in-house accelerators designed with merchant silicon partners, splitting the accelerator market into a training near-monopoly and a contested, price-competitive inference tier.

forming · confidence 40 · Emerging (watchlist) · tracking since August 27, 2026 · updated August 27, 2026

Sign in to get threshold and movement alerts for this trend.

Score history

Daily conviction score, 0 to 100. Higher means the thesis is more strongly corroborated.

Aug 27 · 40Aug 28 · 34

Now 40 · -6 since Aug 27 · ranged 34 to 40

Why the conviction moved

  • Aug 28
    Weakened

    Amazon, the operator furthest along on in-house accelerators with Trainium, committed to two million additional Nvidia GPUs for 2027 and 2028 and put Annapurna Labs to work as Nvidia's first NVHBM memory partner rather than purely on displacement silicon. The largest custom-silicon builder deepening both its purchase volume and its engineering entanglement with the merchant vendor weakens the claim that steady-state inference is migrating off merchant GPUs.

  • Aug 27
    Strengthened +5

    OpenAI's Jalapeño claims 1.5–1.9x throughput per kilowatt versus Nvidia's GB300 on a suite OpenAI chose and ran itself. If the perf-per-watt gap holds under independent testing, the highest-volume steady-state inference workload at one of the largest buyers leaves the merchant GPU tier.

  • Aug 27
    Strengthened +6

    Nvidia is opening NVHBM — a custom memory design moving the controller into the HBM stack for a claimed 30 percent more bandwidth and 15 percent lower memory power than HBM4E — to rival accelerators via NVLink Fusion, with Amazon's Annapurna Labs as the first partner. Nvidia selling its memory advantage to the in-house silicon team most likely to displace its GPUs is an incumbent conceding that the inference tier will be contested and choosing to be paid inside it.

Source trail

  • Supporting · August 27, 2026

    Nvidia Opens Its Custom Memory Design to Rival Accelerators Through NVLink Fusion

    Nvidia is opening NVHBM — a custom memory design moving the controller into the HBM stack for a claimed 30 percent more bandwidth and 15 percent lower memory power than HBM4E — to rival accelerators via NVLink Fusion, with Amazon's Annapurna Labs as the first partner. Nvidia selling its memory advantage to the in-house silicon team most likely to displace its GPUs is an incumbent conceding that the inference tier will be contested and choosing to be paid inside it.

    NVIDIA Blog
  • Supporting · August 27, 2026

    OpenAI Publishes First Benchmarks for Its Jalapeño Inference Chip Against Nvidia Systems

    OpenAI's Jalapeño claims 1.5–1.9x throughput per kilowatt versus Nvidia's GB300 on a suite OpenAI chose and ran itself. If the perf-per-watt gap holds under independent testing, the highest-volume steady-state inference workload at one of the largest buyers leaves the merchant GPU tier.

    AI Post (Telegram)

Unlock full source trail, score history, and daily updates.

Unlock Trends

Affected regions & assets

Assets3 assetsUnlock Trends

Townsquare

Argue the thesis in Townsquare.