SHAPrune: Shapley-Supervised Visual-Token Pruning for Frozen Vision–Language Models
Abstract
We hypothesize that a frozen vision–language model's own removal effects identify the visual tokens it needs better than the attention it assigns them. To test this, we compute offline the Shapley values of the 64 3×3 visual-token blocks of LLaVA-1.5-7B on its own first answer token for 1,400 VQAv2 pairs, and amortize them into 256 weights, one per head of the first eight LLM layers, over each token's text-to-image attention. After layer 8 the scores select tokens by top-k (SHAPrune) or serve as the relevance of a determinantal point process (DPP; SHAPrune-D); inference evaluates no coalition. With depth, features and top-k fixed, the readout keeps 95.9/84.6% of POPE/GQA answers at 64 tokens, against 83.2/70.0% for FastV's score and 91.4/78.3% for the best label-free readout, and a registered fresh-image test confirms the gap to FastV's score (85.3 against 71.4%). As DPP relevance it keeps more answers than CLIP relevance and than none on 10,000 fresh images (registered; 88.01 against 86.58 and 85.45%; accuracy against CLIP relevance: no difference detected). Its 256 head-level weights keep more answers than a per-layer weighting trained on the same labels (registered), and cheaper removal-effect labels, down to single ablations, train heads from which no difference is detected. On 22 of 24 benchmark cells no difference from CLIP relevance at the same depth is detected, and at equal kept tokens a substantial part of SHAPrune-D's margin over earlier diversity-based pruners comes from pruning after layer 8. The model's own removal effects are thus a better pruning signal than its attention scores.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.