Semantics from Structure: Unsupervised Representation Learning from Multi-View Images via Distillation with Feature Fields
Abstract
Unsupervised representation learning typically relies on heavy data augmentation and massive datasets. Reducing these dependencies is necessary for applications in domains where large-scale data collection is impractical. To this end, we introduce Semantics from Structure, a novel paradigm for unsupervised representation learning from multi-view observations that leverages 3D geometry to learn robust visual representations from limited data. Central to our methodology is Distillation with Feature Fields, which performs bidirectional distillation between a radiance field and a randomly initialized image encoder. We demonstrate that our framework successfully learns multi-view consistent semantic representations from a small collection of images without any augmentation or momentum encoders. The learned representations achieve strong downstream performance on 2D and 3D multi-view segmentation across a variety of scenes with no additional fine-tuning. Our approach demonstrates that learning from structural priors yields robust representations with drastically reduced data requirements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.