# A New Paper Names the Networking Bottleneck No Disaggregated Inference System Solves Correctly

When the prefill and decode stages run on separate graphics-processing-unit (GPU) pools, moving the key-value cache between them becomes a data-center data-movement problem, and the authors argue current systems handle it incorrectly.

- Published: 2026-08-03T05:38:34.576Z
- Canonical: https://polylog.news/ai/2026-08-03/a-new-paper-names-the-networking-bottleneck-no-disaggregated
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.LG](https://arxiv.org/abs/2607.28633), [Anthropic Research](https://www.anthropic.com/research/team/economic-research)

Disaggregated inference, which splits the compute-bound prefill stage and the memory-bound decode stage onto separate accelerator pools, has become standard practice at serving scale because it lets each stage be provisioned and batched ind…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-03/a-new-paper-names-the-networking-bottleneck-no-disaggregated (subscription information: https://polylog.news/pricing).