VINO: Revisiting 3D Representation Learning with Volume Transformers
Abstract
Self-supervised representation learning has proven highly effective in 2D, but comparable success has yet to be achieved for 3D scenes. Motivated by this, we identify key limitations in current approaches to 3D representation learning and revisit their design choices. We first replace the commonly used hierarchical PTv3 backbone with the Volume Transformer, a ViT-like architecture for 3D scenes. We then redesign two core objectives of modern approaches built around teacher–student networks: matching representations across views and predicting masked content from visible context. For the former, we separate the roles of teacher and student: the teacher provides clean, consistent targets while the student learns invariance to stronger perturbations. For the latter, we prevent the student from accessing the masked geometry and instead require it to infer missing representations from visible context only. Finally, we redesign 2D-to-3D knowledge distillation by replacing direct regression of view-dependent image features with more view-consistent supervision. Together, these changes form VINO, a representation learning framework that learns semantic representations of 3D scenes without labels. VINO substantially improves dense linear probing across indoor and outdoor benchmarks and transfers strongly to out-of-domain settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.