G2P-WAM: From Geometry Distillation to Preference Alignment for World-Action Models
Abstract
World-action models (WAMs) jointly predict future observations and robot actions, yet visually plausible predictions do not guarantee successful execution. Geometric distillation improves representation grounding, but using geometry for subsequent policy improvement requires relating geometric scores to task outcomes. We propose G2P-WAM, a post-training framework that carries geometric supervision from representation learning into preference alignment. GeoSFT first aligns WAM features with static 3D representations and dynamic 4D trajectories. Reward diagnosis then evaluates candidate geometric scores against execution outcomes and compares their behavior on imagined and observed videos. GeoDPO uses outcomes to order preferences and reconstruction confidence to select trajectories for reference-anchored optimization. Experiments on simulation benchmarks and two physical-robot platforms show that preference alignment further improves the geometrically distilled policy. Joint static and dynamic supervision outperforms either branch alone, while confidence filtering improves on outcome-only preference training with matched pair counts and optimization settings. These results support using geometry for both representation learning and preference selection, without adding geometric teachers or reward scorers at deployment. Project page: https://anonymous.4open.science/w/G2P-WAM-F7BF/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.