acceptodds
Under review as a conference paper at ICLR 2027

TokenButler: Token Importance is Predictable

Abstract

Large Language Models (LLMs) rely on the Key-Value (KV) Cache to store token history, enabling efficient decoding of tokens. As the KV-Cache grows, reading it at every decoding step becomes a major memory-bandwidth bottleneck. However, only a small subset of tokens contribute meaningfully to each decoding step, which is an opportunity to alleviate this bottleneck. A key challenge in finding these *critical tokens* is that they are dynamic, and heavily input query-dependent. Existing methods evict tokens permanently, risking quality; retrieve coarse pages, failing on dense, context-rich tasks; or score every token in the full head dimension, at prohibitive cost. Additionally, many existing KV-Cache sparsity methods rely on inaccurate proxies for token importance. To address these limitations, we introduce **TokenButler**, a high-granularity, query-aware predictor that learns to identify these critical tokens. TokenButler predicts low-dimensional *importance queries* at a fixed depth stride, and combines them with a learned projection of the *real KV-cache keys* to score tokens cheaply, enabling dynamic per-token selection under a fixed budget while preserving the full KV cache. We train TokenButler by distilling the model's masked causal attention distributions, optimizing a lightweight predictor with minimal parameter overhead. On a novel synthetic small-context co-referential retrieval task, designed as a targeted stress test of eviction- and page-based methods, TokenButler achieves near-oracle accuracy where existing methods fail. On general-purpose long-context benchmarks (RULER, LongBench, LongMemEval), TokenButler achieves competitive or superior performance, up to on-GPU speedup using our proposed *prediction interval with neighbor fetching* that amortizes predictor cost while maintaining accuracy within , and up to reduction in latency compared to Dense Attention with CPU offloading, enabling efficient decoding for long contexts in limited GPU memory systems. Code [is available](https://anonymous.4open.science/r/TokenButler-0B6F).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.