ReWorld: Adaptive Corrective Simulation from World Models for Embodied Policy Learning
Abstract
Robotic policies trained from successful demonstrations receive little supervision for recovering from off-nominal states, allowing execution errors to compound into task failure. Physical corrective collection is costly, and task-specific simulators may be unavailable. We introduce ReWorld, which uses an action-conditioned world model as a local simulator to synthesize adaptive corrective experience around demonstrations. ReWorld probes the base policy at predicted perturbation endpoints and retains regions from which it repeatedly fails to recover. For retained cases, a motion planner returns the robot toward the source demonstration and the expert suffix completes the corrective action chunk. The predicted observation and constructed action form corrective training pairs, while competence-aware allocation concentrates supervision on weaker tasks. The world model and planner are used only offline, leaving deployment unchanged. Across 50 RoboTwin 2.0 tasks, ReWorld improves SmolVLA from to and from to . Controlled studies show that selected corrections are more data-efficient than unfiltered corrections, local recovery probes track execution difficulty, and recovery transfers to held-out perturbation magnitudes. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.