# Two New Papers Attack the Least Glamorous Bottlenecks in Serving Small Models

One reformulates the output projection as a vector search to relieve memory bandwidth pressure, the other quantizes the recurrent state that linear-attention models use instead of a growing cache.

- Published: 2026-08-31T06:23:05.714Z
- Canonical: https://polylog.news/ai/2026-08-31/two-new-papers-attack-the-least-glamorous-bottlenecks-in-ser
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2608.27460), [arXiv cs.LG](https://arxiv.org/abs/2608.27513)

Two arXiv papers posted today target costs that only show up once a model is actually being served at volume. The first addresses a bottleneck that grows worse as models get smaller. In a compact multilingual model, the output embedding mat…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-31/two-new-papers-attack-the-least-glamorous-bottlenecks-in-ser (subscription information: https://polylog.news/pricing).