Sub-Patch Semantics in DINOv3: Geometry, Recoverability, and Feature Upsampling
Abstract
Foundational models pretrained on large pool of unlabelled data has brought strong image representation across a wide range of domains. However, their semantically rich but low spatial resolution poses a challenge when working with pixel level dense prediction task such as segmentation or depth prediction. Existing work such as JAFAR, AnyUp try to recover this sub-patch information using the input image as a guide to upsample the ViT feature map. The upsamplers are trained to produce feature map are closer to the feature map as generated from a higher resolution image under the mean absolute error objective. However, a systematic study of what sub-patch information is already encoded within a patch token, how this information is geometrically organized, and to what extent it can be recovered directly from the coarser feature representation remains largely unexplored. In this work, we carry out a study of the geometric and semantic nature of the patch vectors. We find that patch embeddings are anisotropic, with distinct object representations occupying a shared, cone-like subspace where recovering the identity of objects in a mixed patch is challenging. We show that when different objects are present in the same patch, the resulting patch vector can't be explained as a linear decomposition of the mixed tokens. We show that the identity of the foreground object in such cases is still recoverable in a low-dimensional subspace. A Fisher-style criterion identifies such a subspace by selecting directions that separate foreground classes while remaining relatively stable across changes in background. Based on our observation, we propose an upsampler pre-trained on a labelled dataset that learns class prototypes and uses them to extract class-discriminative sub-patch information for coarse DINOv3 patch tokens without requiring an image-based guide that are used by models such as JAFAR or AnyUp. On semantic segmentation, our method improves linear-probe performance by approximately mIoU over the raw DINOv3 tokens, by mIoU over existing image-guided learned feature-upsampling baselines and outperforms transformer and DPT-based decoders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.