Lifetime: KV Cache Eviction from Attention Trajectories for Long-Context Language Models
Abstract
Long-context language models store a key–value (KV) cache whose size grows linearly with the sequence length and quickly dominates inference memory and bandwidth. Training-free prefill eviction reduces this cost, but observation-window scorers typically collapse prompt-tail attention into a single static statistic, discarding its temporal order and conflating tokens whose attention is fading with tokens whose attention is recurring or rising. We introduce , a query-aware, prefill-only method that keeps a temporally binned, ordered attention trajectory for every KV head. Using only evidence available once prefill is complete, estimates a proxy for the proximity of each token's next important use, weights its peak attention rank by this proxy, stabilizes the result with mean-attention rank, and adds a lightweight value-magnitude cue. Every layer–head pair then retains the same number of independently selected entries, which are physically gathered into a regular shortened cache and decoded without rescoring or further eviction. On 16 LongBench tasks, attains the highest average among compressed methods at both and retention on Qwen3-4B ( and ) and Llama-3.1-8B-Instruct ( and ). On RULER-16K with 256 retained entries per layer–KV-head pair, it exceeds the strongest alternative by and points on the two models, while adding less than prefill latency over SnapKV. These results show that prompt-tail temporal structure is a practical, training-free signal for aggressive one-shot KV-cache compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.