acceptodds
Under review as a conference paper at ICLR 2027

Canonical Observation Pretraining for Visuomotor Control

Abstract

Visuomotor policies learned from human demonstrations are brittle to camera viewpoint changes. Existing remedies each make a different assumption: novel-view synthesis requires a rendering pipeline, equivariant architectures require 3D inputs or symmetry hard-coded into the network, and view-invariant pretraining requires only multi-view RGB but suppresses the viewpoint factor, leaving no explicit mechanism for extrapolating viewpoint change. We present , a self-supervised pretraining method that instead keeps viewpoint as a recoverable latent factor, modeling camera viewpoint change as a rotation in SO(3) and warping features into a shared canonical frame before a behavior-cloned policy, which is left unchanged. Pretraining consumes multi-view RGB and the camera angle; at deployment an angle predictor recovers the rotation from the image itself, requiring no pose information. generalizes to camera viewpoints unseen during policy training on both MetaWorld and LIBERO-Goal, outperforming the strongest multi-view baseline by and , and it remains the strongest method at unseen viewpoints on a real-world xArm robot.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.