Human-in-the-World-Model for Scalable Robot Post-Training
Abstract
Post-training robot policies through human correction requires repeated physical execution, scene resets, and operator supervision, limiting the scale of corrective data collection. We propose Human-in-the-World-Model (Hi-WM), a robot policy post-training paradigm that collects human corrective supervision inside an action-conditioned world model, without requiring execution on the target robot during collection. A pretrained policy runs in closed loop on model-generated observations, and a human intervenes when the rollout becomes failure-prone. State caching, rollback, and branching allow operators to revisit the same failure state and collect multiple corrective continuations without repeating physical execution. The generated observations and corrective actions are combined with real-world demonstrations to post-train the policy. Across three real-world manipulation tasks involving rigid and deformable objects and two policy backbones, Hi-WM improves success rates by an average of 37.9 percentage points over the base policies and 19.0 percentage points over a baseline that augments training with successful autonomous world-model rollouts. Increasing the amount of virtual corrective data further improves real-world success in the tested range. These results demonstrate that human corrections collected in a learned world model can improve physical robot policies, supporting world models as reusable environments for corrective post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.