Leveraging Geometric Evidence across Views for Scene Reconstruction
Abstract
Cross-view correspondence provides geometric constraints for multi-view reconstruction by linking observations of the same physical points across views. Recent feed-forward models learn these associations through multi-view feature interactions, often using frame-wise and global attention. However, when matching supervision is confined to a dedicated prediction head, it shapes internal representations only indirectly. Cross-view interactions may therefore remain weakly constrained, limiting how effectively matching information supports reconstruction. We introduce LEVER, which injects cross-view geometric evidence at three levels: output supervision, intermediate feature interactions, and global scene representation. A dense Match Head learns pixel-level cross-view displacement fields, making fine-grained alignment an explicit learning objective. Within the backbone, selected cross-view attention responses are constrained by patch correspondences to encourage information exchange between geometrically related regions. Sequence-level scene tokens aggregate cross-view information under a correspondence-consistency constraint and provide global context to the prediction heads. Together, these components make correspondence a structural signal for reconstruction rather than an auxiliary output. Across five benchmarks, LEVER improves camera-pose estimation, depth estimation, and point-cloud reconstruction over its underlying baseline. Component ablations show that all three levels contribute to reconstruction quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.