acceptodds
Under review as a conference paper at ICLR 2027

Where You Score Matters More Than What You Score

Abstract

Token reduction in vision transformers decides which tokens to keep, merge, or drop. Most existing methods rank tokens by a score computed once, and assume the ranking is independent of the reference set. But the same scoring rule produces different rankings when the reference set changes. We call this the reference-set shift. Changing the reference changes the outcome: relative to full-reference scoring, scoring on the candidate set improves Top-1 on ImageNet-1K by points, with a mean normalized loss improvement of . Because the task loss is a set function, a token's marginal effect depends on interactions with tokens that will be removed, and this dependence is not recoverable from the token's own features. Recovering it requires re-evaluating the candidate set, and matching-based methods do exactly that: they re-evaluate as the set shrinks, so they succeed where scorers computed once cannot. The reference set is therefore an integral part of the method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.