Limitations of relational encoding in visual foundation models
Abstract
Vision Transformers (ViTs) provide dense, patch-level embeddings that support semantic segmentation, depth estimation, and correspondence. Richer visual understanding, however, requires representations to encode physical and spatial relations between entities. We analyze how much of this relational information can be extracted from frozen embeddings of recent foundation ViT models, pre-trained on static images only (DINOv2, DINOv3), on image-text pairs , PE-Core, PE-Spatial), and natively on videos (LeVJEPA). We use single-entity and pairwise probes, validating them on synthetic tasks and applying them to human-object interaction recognition in HICO-DET. To assess whether performance depends on proxy shortcuts, such as pose detection or spatial proximity, we evaluate out-of-distribution (OOD) generalization, position-related shortcuts, and counterfactual inpainting. We find a split between the two probe families. Probes on single entities can detect whether an interaction occurs, but fail to discern which entities actively participate: their accuracy is comparable to that of the best probe trained on positional encoding only. Pairwise probes retain an accuracy gain of points over this positional baseline, and the OOD study shows a degree of generalization across objects. Increasing probe complexity through non-linearity does not yield consistent improvements, with MLP probes performing comparably on average to their linear or bilinear counterparts. Moreover, we find that models exposed to videos or text during training do not perform better on static images. These results suggest that the successes of ViTs in dense semantic extraction from local patches do not extend to the understanding of physical interactions in static images.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.