acceptodds
Under review as a conference paper at ICLR 2027

VINO: Revisiting 3D Representation Learning with Volume Transformers

Abstract

Self-supervised representation learning has proven highly effective in 2D, but comparable success has yet to be achieved for 3D scenes. Motivated by this, we identify key limitations in current approaches to 3D representation learning and revisit their design choices. We first replace the commonly used hierarchical PTv3 backbone with the Volume Transformer, a ViT-like architecture for 3D scenes. We then redesign two core objectives of modern approaches built around teacher–student networks: matching representations across views and predicting masked content from visible context. For the former, we separate the roles of teacher and student: the teacher provides clean, consistent targets while the student learns invariance to stronger perturbations. For the latter, we prevent the student from accessing the masked geometry and instead require it to infer missing representations from visible context only. Finally, we redesign 2D-to-3D knowledge distillation by replacing direct regression of view-dependent image features with more view-consistent supervision. Together, these changes form VINO, a representation learning framework that learns semantic representations of 3D scenes without labels. VINO substantially improves dense linear probing across indoor and outdoor benchmarks and transfers strongly to out-of-domain settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.