BridgeView: Enhancing Cross-View Consistency of Vision Foundation Models via Frozen Opposite-View Targets
Abstract
Vision foundation models provide powerful frozen features, yet image-based pretraining does not require two observations of the same surface to share a descriptor. Existing approaches enforce this relation with an intermediate scene representation and a rendered teacher. Reconstruction is costly, and the rendered teacher is a visibility-weighted mixture in the reconstructed field, so matching it moves student features away from pretrained coordinates. Leaving that coordinate system raises scene-level agreement while frozen semantic transfer falls. Instead, we establish cross-view consistency directly inside the pretrained coordinate system. To this end, we propose BridgeView, which uses the frozen opposite-view descriptor as the teacher on a same-surface pair formed by calibrated geometry. When the encoder cell at that pixel mixes neighboring instances, a bounded residual refines this readout with instance context while remaining inside a cone of the frozen vector. When many pairs observe one instance relation, their backward updates are coordinated inside that relation. We validate BridgeView on indoor RGB-D benchmarks. Across DINOv3 backbones, it raises cross-view correspondence, semantic segmentation, visual localization, and TSDF map rendering, and trains over faster than SnD. Scaling the pair corpus to 3M further improves correspondence. These results suggest a scalable route for large vision foundation models to obtain cross-view consistency inside the pretrained coordinate system. Code is provided in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.