ARGaM: Action-Recoverable Gaussian Motion for Robotic Manipulation
Abstract
Gaussian world models predict future scene evolution in explicit 3D representations, offering geometric cues for robotic manipulation. However, visually accurate predictions alone do not ensure that the underlying motion can reliably guide action trajectories. To address this, we propose ARGaM, which learns Gaussian dynamics autoregressively with temporally aligned visual and action supervision, and refines directly regressed trajectories using a separately recovered geometric motion prior. The dynamics model predicts successive Gaussian states and a coarse action trajectory from shared temporal features, supervised by future-image reconstruction and demonstrated actions across the prediction horizon. An additional motion loss supervises selected gripper Gaussian displacements using demonstration-derived targets, bringing explicit motion constraints into the predicted geometry. Geometric registration then produces a motion prior from the predicted Gaussian correspondences. A residual diffusion model combines this prior, its recovery confidence, the coarse trajectory, and rendered future visual context to predict corrections to the coarse actions. On nine RLBench tasks, ARGaM achieves a mean closed-loop success rate of 72.7%, improving over the strongest evaluated baseline (60.4%) by 12.3 percentage points. On six real-robot tasks, ARGaM succeeds in 48 of 60 trials (80.0%), compared with 35 of 60 (58.3%) for the same baseline, an improvement of 21.7 percentage points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.