# VarRate Cuts Long-Context Memory by Varying KV-Cache Compression Token by Token

The training-free method targets the key-value cache, the dominant memory cost of long-context inference, without the structural limits of token-selection or fixed-rate approaches.

- Published: 2026-07-20T05:31:45.211Z
- Canonical: https://polylog.news/ai/2026-07-20/varrate-cuts-long-context-memory-by-varying-kv-cache-compres
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2607.15498)

A new paper introduces VarRate, a training-free method for variable-rate compression of the key-value cache in long-context large language model (LLM) inference. The key-value (KV) cache stores the attention state for every token in the con…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-07-20/varrate-cuts-long-context-memory-by-varying-kv-cache-compres (subscription information: https://polylog.news/pricing).