What Survives Token Merging? Merge-Invariant Low-Rank Attention for Vision Transformers
Abstract
Token reduction and low-rank factorization offer complementary ways to accelerate vision transformers, yet their composition introduces dependencies that neither addresses independently. Factorizing the joint query–key product changes the representations used for token merging, while merging destroys the stable positional correspondence required by conventional logit-residual correction. We resolve both incompatibilities without retraining. First, we characterize fixed logit corrections invariant to patch relabelling: their space is five-dimensional, but two directions are invisible to the softmax, leaving exactly three effective degrees of freedom. These three coefficients per head recover and of the accuracy recovered by a -entry position-indexed correction on DeiT-Small and DeiT-Tiny, respectively, while remaining applicable after merging. Second, we decouple token matching from the current query–key factorization by computing similarity from normalized block inputs, improving matching stability under rank reduction. We further identify an accuracy–compute crossover: mild token reduction favors merging alone, whereas more aggressive reduction benefits from combining both compression axes. Beyond this crossover, our method improves accuracy by up to percentage points at matched compute and achieves up to throughput at matched accuracy, with no retraining or gradient computation. Experiments across DeiT scales, ViT-L, and an alternative token-reduction rule establish that the composition holds beyond a single setting, and results on ADE20K segmentation and five CLIP image encoders examine the broader applicability of low-rank correction. Code and compressed checkpoints will be released to support reproduction and further work.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.