acceptodds
Under review as a conference paper at ICLR 2027

Bridges over Heavy Hitters: A Causal-HITS Theory and Counterfactual Benchmark for Attention-Graph KV-Cache Compression

Abstract

Training-free KV-cache compression decides which past tokens a long-context language model may forget. Heavy-hitter methods retain tokens with large accumulated attention, whereas Stream-Influence scores tokens by centrality in the sparse causal attention graph. We show that Stream-Influence is a one-pass causal form of Kleinberg's HITS: a token's authority is the attention it receives from later tokens, weighted by their hub scores. Without normalisation, attention sinks make the recursion diverge, and we give a bounded alternative. On bridge-chain graphs, sparsified centrality recovers every bridge once the budget exceeds the chain length, while accumulated attention keeps only a prefix proportional to the budget, because causal mass makes early tokens heavy hitters regardless of content. A symmetry theorem then shows that if several structurally identical chains are present and only the question identifies the gold chain, every score depending on the question through a single hop assigns the same value to a gold bridge and its twin. This class includes H2O, SnapKV, StreamingLLM and Stream-Influence. BridgeBench instantiates both regimes with known bridges. Physical eviction on models from 3B to 14B confirms the obstruction: unseeded scores retain at most 8% of full-cache accuracy at a 10% budget and 21% at 20%. A beam-restricted walk seeded at the question over coreference heads recovers the gold chain, retaining 55% and 81% at these budgets against 10% and 25% for the best baseline. The same heads, transferred without re-calibration, keep 74% of the gold sentences on HotpotQA, where SnapKV keeps 12%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.