GradSelect: Output-Targeted Patch Ranking under Exact-K Vision Transformer Pruning
Abstract
How should a frozen vision transformer choose K patches when the pruning layer and token budget are fixed? This paper isolates patch ranking from system choices such as pruning schedules, routing modules, and classifier changes. GradSelect scores each image by first running the full model to obtain its own predicted class, backpropagating that output to the cut layer, keeping positive first-order support for each patch, and fusing this signal with CLS attention after per-image z-scoring. The fusion weight is selected on a lock split, and all final comparisons use the same exact-K intervention. At 25% retention after block 8, GradSelect has the highest point estimate among primary selectors in all nine ImageNet-1K, Food-101, and SUN397 cells across CLIP, DINOv2, and DeiT, including five cells not used to choose the recipe. Against the stronger per-cell opponent drawn from either the primary suite or the attribution/removal stress suite, it wins 6/9 cells with +2.24-point mean and +2.53-point median accuracy margins; exact single-token removal remains stronger on supervised DeiT. On Caltech-101, GradSelect is best on CLIP and DINOv2 and 0.22 points below DivPrune on DeiT. Because scoring requires a full forward pass and one backward pass, GradSelect is best understood as an offline output-aware patch selector rather than an online accelerator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.