Beyond View Consistency: Rotation-aware Consistent Alignment for Viewpoint-Robust Robotic Manipulation
Abstract
Visual robotic manipulation policies remain highly sensitive to viewpoint variations, which can significantly degrade performance. Existing approaches commonly address this issue by learning view-consistent representations. However, directly enforcing cross-view consistency may overlook viewpoint-dependent task-relevant information for precise manipulation. To address this limitation, we propose Rotation-aware Consistent Alignment (RCA). RCA prioritizes capturing task-relevant features. RCA then accounts for camera rotation to adapt viewpoint-dependent features and align them in a unified latent space. In practice, our framework comprises two stages. The first stage learns task-relevant representations from each observation. The second stage adapts and aligns these representations across viewpoints using a FiLM mapping network conditioned on predicted viewpoint parameters. We evaluate RCA in Meta-World and PandaGym simulations and real-world robot experiments under diverse viewpoint perturbations. RCA improves average success rates over baselines by 19.8, 26.7, 4.2, and 6.9 percentage points in Meta-World azimuth, Meta-World camera shaking, PandaGym azimuth, and real-world azimuth evaluations, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.