Decoupling Content and Spatial Support in Frozen Vision Language Readouts.pdf
Abstract
Dense prediction with a frozen vision–language encoder depends on both the content supplied to its text-aligned head and the spatial support used to aggregate patches. We investigate whether the same representations are suitable for both roles by generalizing the pretrained pooling head into a support-set decoder with independently selectable content and support depths. An exhaustive 13x13 depth map on SigLIP 2 B/16 reveals distinct preferences: final-layer content and pre-contextual support. Holding content fixed, replacing pre-contextual support with final-layer support reduces VOC20 mIoU by 6.4 points. Statistical controls clarify the distinction: matching content from an earlier layer to the native head input's moments recovers most of its readout loss, while matching final-layer support features to pre-contextual moments substantially narrows the support gap. These interventions distinguish compatibility with the frozen head from the feature geometry used for spatial aggregation. Cross-encoder controls characterize where the shallow-support preference holds. Applying this separation yields SigMAP, a training-free decoder that combines final-layer content with pre-contextual support and improves on identity support across five datasets. Together, these findings identify content compatibility and spatial support quality as distinct criteria for selecting representations under a frozen readout.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.