acceptodds
Under review as a conference paper at ICLR 2027

Learning Geometrically-Grounded Amodal 3D Representations for View-Generalizable Robotic Manipulation

Abstract

Real-world robotic manipulation requires visuomotor policies capable of robust 3D scene reasoning under varying camera viewpoints. While recent 3D-aware manipulation policies have shown promise, they still face several limitations: (i) reliance on multi-view observations during inference, which is impractical in camera-constrained deployments; (ii) insufficient geometric fidelity in learned representations, limiting precise and viewpoint-robust control; and (iii) the lack of effective mechanisms for transferring pretrained 3D representations to downstream visuomotor policies. To address these challenges, we present GEM3D (Geometrically-Grounded 3D Manipulation), a unified 3D representation and policy learning framework for view-generalizable robotic manipulation. GEM3D learns amodal 3D representations from single-view RGB-D observations by jointly enforcing point-cloud reconstruction and Gaussian-splatting-based novel-view consistency during pretraining. A multi-step distillation strategy then transfers the learned geometric understanding into a deployable single-view visuomotor policy for downstream manipulation control. We evaluate 3D scene reasoning through zero-shot viewpoint generalization under unseen camera poses. Extensive experiments across 12 RLBench tasks and three real-world robotic manipulation tasks demonstrate that GEM3D achieves strong viewpoint robustness and significantly outperforms existing state-of-the-art methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.