Old-School Warping Shines for Modern Self-Supervised 3D Reconstruction
Abstract
Humans can develop a rich understanding of the 3D world with little to no explicit geometric supervision. Inspired by this ability, we show that a simple multi-view Transformer can learn strong 3D reconstruction capabilities through self-supervision with an "old-school" image warping loss. Our model, Warp3R, jointly predicts depth maps and camera poses and is trained using photometric and geometric consistency across views. It substantially outperforms state-of-the-art self-supervised methods on camera pose estimation across diverse datasets and improves self-supervised depth estimation. Extending supervised training with our objective on unlabeled data further improves both pose and depth estimation. Our results highlight the effectiveness of classical geometric principles for learning 3D reconstruction with modern networks. Code will be released to facilitate future research
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.