acceptodds
Under review as a conference paper at ICLR 2027

Rollback: Experience Backtracking in World Models for Policy Improvement

Abstract

Action-conditioned world models offer a scalable alternative to costly real-world interaction for improving robot policies. However, existing world-model-based methods largely treat world models as forward-only surrogates of the physical environment, leaving untapped a capability no physical environment affords: rewinding to any past state to explore counterfactual futures. In this work, we introduce Rollback, a policy improvement paradigm that turns world models from passive simulators into counterfactual engines: it rolls back to critical decision points, branches into alternative actions, and contrasts their outcomes to attribute success to individual decisions. This concentrates imagination where it matters most and distills counterfactual outcomes into fine-grained, action-level supervision without a separately learned value model. Furthermore, we build a scalable training system that unifies distributed world-model search with asynchronous real-robot execution, forming a self-reinforcing cycle: real-world experience continually sharpens the world model, and a sharper world model in turn drives further policy improvement. Experiments on three challenging real-world tasks show that Rollback improves success rates over the strongest baseline by 10.44%, 11.11%, and 11.85% on block extraction, ball tossing, and tabletop curling, respectively, while achieving more stable improvement throughout training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.