# Small Models and Compression Research Push More Inference Off the Cloud

A widely circulated developer essay argues small models have become good enough for real products, while new work on binarized network pruning and multi-drafter speculative decoding attacks the cost from both ends.

- Published: 2026-08-28T06:11:21.137Z
- Canonical: https://polylog.news/ai/2026-08-28/small-models-and-compression-research-push-more-inference-of
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Hacker News](https://calv.info/small-models-have-arrived), [arXiv cs.LG](https://arxiv.org/abs/2608.26233), [arXiv cs.CL](https://arxiv.org/abs/2608.26112</source_url_placeholder)

An essay titled "Small Models Have Arrived" reached the top of Hacker News this week, arguing that models small enough to run locally now meet the quality threshold for a large class of production tasks, and that the practical constraint on…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-28/small-models-and-compression-research-push-more-inference-of (subscription information: https://polylog.news/pricing).