acceptodds
Under review as a conference paper at ICLR 2027

LaRA-WM: Test-Time Latent Action Refinement for Robotic Manipulation in Dynamic Environments

Abstract

Robotic manipulation policies trained from visual observations often achieve strong in-distribution performance, yet remain brittle under out-of-distribution changes such as novel objects, altered lighting, and unseen scene layouts. A key limitation is that most existing policies directly map high-dimensional observations to low-level motor commands, leaving little room for deliberation or correction before action execution. In this paper, we propose LaRA-WM (Latent Robot Action World Model), a framework that shifts robotic control from direct action regression to reward-guided optimization in a compact, task-centric latent action space. During offline training, LaRA-WM learns visual and language-conditioned latent action representations from expert demonstrations, together with an action decoder and a reward-feature scorer that evaluates candidate latent intents. At test time, instead of producing a one-shot action, LaRA-WM iteratively refines latent actions using the Cross-Entropy Method guided by the learned scorer, and only then decodes the optimized latent intent into an executable motor command. This design enables the policy to “think before acting” in a structured semantic space, avoiding the instability and local optima associated with raw action-space search. Experiments in the RoboTwin manipulation benchmark show that LaRA-WM consistently improves task success rates, returns, and out-of-distribution generalization over direct policies and non-refinement latent baselines. These results suggest that test-time latent action refinement provides an effective and efficient mechanism for robust robotic manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.