acceptodds
Under review as a conference paper at ICLR 2027

GroundingWAM: Learning Where to Focus in World Action Models

Abstract

World Action Models (WAMs) achieve high success rates on manipulation benchmarks, yet can struggle when execution departs from familiar demonstration trajectories. Following an unexpected object drop, a policy may fail to relocalize the target and resume manipulation. We probe this grounding weakness through auxiliary localization supervision: the predicted target object location can remain near the gripper even after the object has moved elsewhere. We interpret this behavior as a consequence of the strong correlations between object positions, robot states, and task progress in successful demonstrations, which allow models trained with end-to-end regression objectives to exploit predictable patterns without consistently grounding their predictions in current visual evidence. Motivated by this observation, we introduce GroundingWAM, which uses task-relevant spatial and language annotations to directly supervise internal attention during action generation, guiding action queries toward relevant image regions and instruction tokens. Trained on successful demonstrations without additional recovery demonstrations, GroundingWAM achieves 99.3% success on standard LIBERO and 87.1% on LIBERO-Plus benchmark. On our object-release setting, it improves success from 47.79% to 64.71%, with a gain of 41.75 percentage points on LIBERO-Spatial. Component comparisons and rollout visualizations support internal grounding supervision as a practical way to strengthen manipulation beyond familiar execution patterns.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.