Canonical Observation Pretraining for Visuomotor Control
Abstract
Visuomotor policies learned from human demonstrations are brittle to camera viewpoint changes. Existing remedies each make a different assumption: novel-view synthesis requires a rendering pipeline, equivariant architectures require 3D inputs or symmetry hard-coded into the network, and view-invariant pretraining requires only multi-view RGB but suppresses the viewpoint factor, leaving no explicit mechanism for extrapolating viewpoint change. We present , a self-supervised pretraining method that instead keeps viewpoint as a recoverable latent factor, modeling camera viewpoint change as a rotation in SO(3) and warping features into a shared canonical frame before a behavior-cloned policy, which is left unchanged. Pretraining consumes multi-view RGB and the camera angle; at deployment an angle predictor recovers the rotation from the image itself, requiring no pose information. generalizes to camera viewpoints unseen during policy training on both MetaWorld and LIBERO-Goal, outperforming the strongest multi-view baseline by and , and it remains the strongest method at unseen viewpoints on a real-world xArm robot.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.