GrabMe: Accelerating ViTs for Segmentation through Graph-based Token Merging
Abstract
Segmentation is a widely studied computer vision task in which Vision Transformers (ViTs) have demonstrated notable performance. Nevertheless, the self-attention mechanism in ViTs incurs quadratic computational complexity with respect to the number of tokens, which has motivated the development of a range of token-merging strategies. Still, existing merging methods based on global matching or rigid local grouping may perturb critical spatial relationships and degrade the preservation of fine-grained details. Furthermore, reliance solely on pairwise similarity metrics may be inadequate for preserving object boundaries and maintaining the integrity of small objects during the merging process. In this work, we introduce Graph-based Token Merging (GrabMe), a training-free token-merging method for efficient segmentation. To more faithfully preserve relative spatial structure, GrabMe constrains merging to one-hop neighbors on a spatial adjacency graph, where resulting token groups inherit adjacencies of their constituent tokens. For parallel merging, GrabMe employs dynamic programming to construct a conflict-free pool of candidates with maximum cardinality. Subsequently, complementary structural context derived from one-hop similarity distributions is incorporated into the merge decision, thereby mitigating the loss of fine-grained details. Extensive experiments with ViT backbones across semantic, instance, and panoptic segmentation benchmarks demonstrate that GrabMe effectively improves the accuracy–efficiency trade-off of pretrained backbones for segmentation without requiring any additional training. Our code is available at this https://anonymous.4open.science/r/GrabMe-7889/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.