Learning Latent 3D Rotations from Unlabeled Visual Sequences
Abstract
Learning how visual states transform under physical transformations is important for visual reasoning, but transformation-aware models typically assume access to transformation labels. We ask whether -parameterized actions, learned from unlabeled visual sequences, can recover meaningful structure of the underlying physical 3D rotations. Given object-centric sequences of rotating objects, we jointly train an action inference network and an action-conditioned transition model in the latent space of a frozen visual encoder. To discourage degenerate predictive codes, we introduce a reverse-prediction constraint that requires inferred actions to support consistent inverse transitions. Without using labeled data, the learned actions recover meaningful structure of the underlying 3D rotations. As a complementary evaluation, we show that the learned dynamics support object matching across 3D viewpoint changes through latent-action planning, outperforming static feature-similarity baselines on abstract shapes with limited semantic shortcuts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.