acceptodds
Under review as a conference paper at ICLR 2027

PlayWAM: Learning Dynamics and Value in World Action Models through Controlled Counterfactual Play

Abstract

World Action Models (WAMs) serve as both dynamics models and action generators. However, this dual role creates conflicting data requirements: dynamics learning benefits from diverse experiences beyond successful execution, whereas action generation requires supervision selected for task success. Our key insight is that simulation enables controlled alternative futures from shared task states, revealing both physical consequences and their task utility. Motivated by this, we introduce PlayWAM, which combines branch-specific learning with controlled counterfactual play—a simulation-native pipeline for generating comparable alternative futures from shared states. We use state cloning and parallel rollouts with controlled interventions to capture nominal, perturbed, failing, and recovering behaviors. PlayWAM learns dynamics from all recorded branches but imitates only task-successful segments; other observed actions condition future prediction rather than serve as imitation targets. A lightweight value head ranks same-state branches and scores parallel future–action candidates, enabling single-round selection without outer-loop refinement. We evaluate on PlayBench, RoboTwin-Play, LIBERO, and eight real-world tasks. Project page: https://stellar-wisp-d1ecc3.netlify.app/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.