GuardWAM: Mode-Conditioned Structural Priors for Contact-Rich Video World Models
Abstract
Video world action models (WAMs) predict future video conditioned on robot actions, yet visual plausibility does not ensure action faithfulness: when contact changes, imagined motion may violate what the action can physically cause. Using physics-grounded paired probes, we find that a frozen LingBot-VA checkpoint localizes action effects significantly worse at contact onset than in matched stable controls, despite unchanged global visual quality. Neither a mode token nor global identity/inverse/composition regularization repairs this failure; the latter improves reversible free-space behavior while inducing false restoration at irreversible transitions. This reveals a missing principle: structural relations should hold only within interaction-dependent validity domains. We instantiate this principle in GuardWAM. An interaction guard predicts mode, transition events, and executability; a deterministic validity compiler activates only valid structural constraints; and mode-flow plus boundary jump/reset adapters model within-mode evolution and discontinuous transitions. On held-out RoboTwin probes, Oracle GuardWAM reduces contact-boundary false restoration and false success to roughly one third of the frozen WAM while preserving free-space consistency. A learned guard retains about 90% of this repair, identifying guard accuracy as the remaining bottleneck.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.