acceptodds
Under review as a conference paper at ICLR 2027

Beyond Token Content: Visual Token Retirement by Tracking Evolving Influence

Abstract

Token compression inside the vision encoder can reduce computation in both subsequent vision blocks and the downstream language model, but a key challenge is determining when a token can be safely retired while its representation is still evolving. Existing criteria largely rank tokens by redundancy estimated from their current representations or local relations, using signals such as salience and representativeness, without directly assessing whether subsequent computation still depends on each token's continued participation. Through controlled interventions, we show that retaining an intermediate token state as frozen key–value memory does not substitute for allowing the token to continue updating and interacting across later blocks. We therefore recast encoder-internal compression as token retirement and propose Influence-Confirmed Token Retirement (ICTR), a training-free and instruction-agnostic procedure that progressively shortens the sequence to a fixed terminal budget. We estimate external influence without additional forward passes by analytically computing how removing a token from the current key–value set would alter other tokens' outputs and averaging the largest relative changes. As representations and the active set evolve, we recompute external influence and confirm candidates whose ranks remain low across consecutive blocks. Across 10 benchmarks with LLaVA-1.5-7B, the method retains over 96% of dense-model performance on average at the 64-token budget, after retiring 88.9% of visual tokens; at the 128-token budget, it achieves a prefill speedup and reduces the KV-cache footprint by 71.2%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.