Spatial Mixing Predicts the Read-out Depth for Geometry in Vision Foundation Models
Abstract
Dense geometric tasks are read off a vision transformer's output block by convention, under the suspicion that depth trades spatial detail for abstraction. We find instead that the depth at which a linear probe best decodes a geometric attribute rises with the attribute's ground scale, so the best depth belongs to the attribute and not to the network alone. Over 32 backbones this ordering appears in every one whose coarse-scale representation is still improving at the output block, and it transfers to indoor photographs. Our account is that each block mixes a token with its neighbours over a length that grows with depth, so depth loses geometry below that length and keeps what is coarser. We measure that length without labels, as the radius over which a token responds to one occluded input patch. On overhead imagery, where every token has a known ground scale, dividing each target's scale by that radius brings the retention curves of nine matched backbones together for surface normals and aspect and not for band-passed height, and no rival length of the mixing does this. Blurring an intermediate block reproduces the output block's loss at a width the radius predicts. Narrowing late attention as far as semantic accuracy allows returns the lost surface geometry on three of six backbones and worsens it on none. Mixing is therefore part of the cause. An earlier block still holds the lost detail, and on the same imagery warping a tile by a known map names it without labels or training, recovering 80 to 88% of what a labelled sweep would gain over all 32 backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.