Rotation Sign in Frozen Vision Encoders: What Token Readouts Reveal
Abstract
A global image embedding can obscure information that remains accessible in the encoder's tokens. We study this distinction through rotation-sign prediction: identifying whether an image was rotated clockwise or counterclockwise. Across layers of frozen vision encoders, we compare readouts that retain token position with readouts that aggregate over it. In CLIP-B/32, a linear readout of final-layer token positions reaches 92.4% accuracy, whereas global averaging and the tested nonlinear permutation-invariant readouts remain near the 50% chance level. Training on shuffled token positions reduces accuracy to 51.2%, supporting a role for spatial organization in this contrast. DINOv2-B behaves differently: both pooled features and nonlinear invariant readouts retain substantial rotation-sign information. Additional experiments show that padding can introduce content-free cues and that probes trained on synthetic transformations do not yield a consistent advantage over permutation controls on a spatial-relation task. These results identify an encoder-dependent gap between token-level and global readouts. They also show why poor probe performance on an embedding should not be taken as evidence that the encoder has discarded the corresponding information. Establishing whether the gap is caused specifically by aggregation requires tighter capacity and optimization controls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.