Scr4tch: Self-Supervised 4D Reconstruction from Raw Video Alone
Abstract
Self-supervised learning has enabled powerful visual representations at scale, yet its potential for understanding the 3D structure and dynamics of the world directly from raw video remains largely unrealized. We present Scr4tch, a model that learns geometrically grounded 4D representations entirely from scratch using uncalibrated monocular videos. Without requiring camera calibration, geometric annotations, or pretrained geometric models, Scr4tch unlocks self-supervised learning for 4D understanding from large-scale Internet videos. Our key idea is to learn 4D representations that encode cross-view relationships through latent reconstruction, while constraining how motion is modeled to separate camera motion from scene dynamics and suppress appearance shortcuts. Joint geometric optimization then grounds these representations in an explicit, deformable 4D scene. To train Scr4tch, we additionally collect 2,154 hours of challenging, casually captured videos from YouTube without offline geometric preprocessing. Experiments demonstrate accurate camera estimation, 4D reconstruction, and high-quality novel view synthesis on unseen videos, establishing a path toward scalable self-supervised learning of the dynamic world.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.