# Two Speculative-Decoding Papers Attack the Fragility That Limits Inference Speedups

One splits drafting between an on-device small model and a server verifier across different vocabularies, the other pretrains drafters without a fixed target model to stop acceptance rates collapsing under workload shift.

- Published: 2026-09-10T06:24:02.019Z
- Canonical: https://polylog.news/ai/2026-09-10/two-speculative-decoding-papers-attack-the-fragility-that-li
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2609.09166), [arXiv cs.CL](https://arxiv.org/abs/2609.09338), [NVIDIA Blog](https://blogs.nvidia.com/blog/ibc-news-2026/)

Speculative decoding, in which a small model drafts tokens that a large model verifies in parallel, is the standard way to cut latency for reasoning models. Its weakness is well known to anyone running it in production: drafters are trained…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-10/two-speculative-decoding-papers-attack-the-fragility-that-li (subscription information: https://polylog.news/pricing).