acceptodds
Under review as a conference paper at ICLR 2027

ReWorld: Adaptive Corrective Simulation from World Models for Embodied Policy Learning

Abstract

Robotic policies trained from successful demonstrations receive little supervision for recovering from off-nominal states, allowing execution errors to compound into task failure. Physical corrective collection is costly, and task-specific simulators may be unavailable. We introduce ReWorld, which uses an action-conditioned world model as a local simulator to synthesize adaptive corrective experience around demonstrations. ReWorld probes the base policy at predicted perturbation endpoints and retains regions from which it repeatedly fails to recover. For retained cases, a motion planner returns the robot toward the source demonstration and the expert suffix completes the corrective action chunk. The predicted observation and constructed action form corrective training pairs, while competence-aware allocation concentrates supervision on weaker tasks. The world model and planner are used only offline, leaving deployment unchanged. Across 50 RoboTwin 2.0 tasks, ReWorld improves SmolVLA from to and from to . Controlled studies show that selected corrections are more data-efficient than unfiltered corrections, local recovery probes track execution difficulty, and recovery transfers to held-out perturbation magnitudes. Code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.