GeoIDM: Action-Reprojected Geometric Supervision for Generalizable Inverse Dynamics
Abstract
Video-based robot policies can decouple visual planning from control by translating predicted future observations into actions with an inverse dynamics model (IDM). However, IDMs supervised solely by absolute end-effector poses and gripper states can overfit appearance cues and generalize poorly to unseen trajectories and tasks. We present GeoIDM, a Geometry-supervised IDM that grounds action prediction in temporally aligned visual correspondences. A frozen VGGT encoder extracts visual tokens with geometric information from current and future observations, and a lightweight query compressor conditions an action DiT. During training, GeoIDM applies action-reprojected geometric supervision by accumulating predicted translations to estimate the future end-effector position, reprojecting it into the future view using camera calibration, and aligning the sampled feature with the corresponding feature in the current frame through a local contrastive objective. A gate suppresses static or invalid projections. The geometric branch is removed at inference, so the model still maps visual observations to actions without an additional geometry module. We evaluate GeoIDM through offline action prediction on AgiBot, replay of reference videos in LIBERO and RoboTwin 2.0, and replay of opening, clicking, and grasping tasks on a physical Franka robot. Relative to the DreamGen baseline, GeoIDM reduces position error by 63.9% and 29.3% on the cross-episode and cross-task AgiBot splits, respectively. It further improves cross-task replay success by 20.5 and 18.0 percentage points on LIBERO and RoboTwin 2.0, respectively. These results show that geometric supervision during training improves IDM generalization while preserving the standard interface at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.