Polylog
The Polylog AI Intelligence Brief

Morning Edition · Monday, July 20, 2026Published at 1:31 AM EDT · New York

VarRate Cuts Long-Context Memory by Varying KV-Cache Compression Token by Token

The training-free method targets the key-value cache, the dominant memory cost of long-context inference, without the structural limits of token-selection or fixed-rate approaches.

VarRate Cuts Long-Context Memory by Varying KV-Cache Compression Token by Token

A new paper introduces VarRate, a training-free method for variable-rate compression of the key-value cache in long-context large language model (LLM) inference. The key-value (KV) cache stores the attention state for every token in the con…

Continue the AI Intelligence Brief

Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.

  • 5 AI intelligence signals a day
  • Frontier labs, compute, and chips
  • Model releases and AI infrastructure
  • Source-grounded analysis with confidence labels

The Global Intelligence Brief stays free.

Part of a tracked trend

The Inference-Cost Efficiency Race

Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.