acceptodds
Under review as a conference paper at ICLR 2027

Below the Patch: Separating Information from Accessibility in Frozen Vision Transformers

Abstract

Vision transformers represent an image on a coarse patch grid, yet their frozen features are widely used for tasks requiring localization far below the patch stride. What, then, does “resolution” mean for a frozen representation? We argue that answering this question requires separating information from accessibility: a signal may be present in the input or representation while remaining inaccessible to a particular comparison and readout. We study this distinction using controlled sub-pixel displacement. Raw patches contain enough evidence to recover 2-D motion when local image structure is well conditioned ( up to ), even though a linear probe on the same pixels reports essentially zero. The ViT patch embedding is not the bottleneck either: on our grayscale input subspace it is injective, and a nonlinear decoder recovers displacement equally well from raw patch pairs and layer-0 token pairs (). The apparent loss emerges when the two observations are collapsed into an aligned difference, where pixel and token inputs fail alike. This bottleneck is structured rather than complete. The aligned difference retains the aperture-constrained component of motion in a linearly accessible form. Tracking this coordinate through depth reveals that frozen checkpoints differ sharply in how much fine-scale geometry they keep accessible: MAE largely preserves it, whereas a label-supervised ViT reduces late-layer linear accessibility to nearly zero. Intervening on the fitted direction strongly controls the downstream readout, and combining the accessible scalar with the corresponding image geometry supports sub-pixel refinement on naturalistic video. These results show that effective spatial resolution is not determined by patch stride alone. It depends on what information is present, how observations are compared, and what the readout can access.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.