When Less Retrieves More: Selective Token Retention in Visual Retrieval
Abstract
Can discarding visual patch vectors produce a better retrieval representation than aggregating them? We study this question with frozen encoders, full query tokens, and exact MaxSim scoring, developing the study on revisited Oxford. Retaining 39 of 196 database patches by CLS attention improves DINOv2 Medium mAP from 63.91% to 72.30%, with 80.1% fewer stored patch vectors. Holding those attention-selected anchors fixed, averaging all patches into their nearest anchors reduces mAP to 69.13%; normalizing those means reduces it further to 64.57%. Retention also outperforms nine specified hierarchical-pooling adaptations at the same exact count. An additional Paris evaluation improves Medium mAP from 85.92% to 88.23% and Hard mAP by 4.20 percentage points relative to the full representation. The attention-versus-full gain also recurs with DINOv3 on Oxford, but dense CLIP changes by -1.61 points; native DINOv3 CLS remains a strong single-vector alternative. Ordering diagnostics show asymmetric repair and damage, while CLIP demonstrates that improved mean margins against fixed competitive negatives do not suffice to explain improved mAP. The contribution is a controlled selection-aggregation contrast in DINOv2 patch retrieval, not a new compression algorithm or a universal benefit of deletion. The Paris evaluation extends the direction of the attention-versus-full gain, not the Oxford-only same-anchor finding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.