Dream to Remember: Continual Steering of Generalist Robot Policies with World Models
Abstract
Generalist robot policies are trained with imitation learning and are bounded by the quality and coverage of their demonstrations. Reinforcement learning can improve existing skills and add new ones, either by fine-tuning the policy weights or by training a small module to steer the input noise of a frozen generative policy. In open-ended deployment, tasks arrive sequentially, and both approaches suffer from catastrophic forgetting. Fine-tuning on a task degrades performance on prior tasks, and steering only shifts forgetting to the shared steering module. We propose Dream to Remember (D2R), which continually steers a frozen generalist policy toward new tasks, one after another, while retaining its earlier skills. As the steering module learns a new task, a frozen pretrained world model generates imagined rollouts of prior tasks, anchoring the module to its earlier behavior. Prior tasks are thus rehearsed without a replay buffer or new data. When the world model is faithful enough, D2R learns the new task entirely in imagination. Otherwise, it learns the task in the environment while using the world model only to rehearse prior tasks. On CALVIN, D2R improves eight tasks in sequence entirely in imagination, raising average success from 63.7% to 74.8%, while LoRA fine-tuning on the same sequence drops it to 11.7%. With a frozen off-the-shelf generalist policy and world model, and A2World, D2R learns unseen LIBERO-90 tasks in the environment while keeping 's performance on the LIBERO suites it was fine-tuned on, even though the world model is not accurate enough for policy optimization. Finally, on a real Franka robot, D2R improves three tasks in sequence entirely in imagination while retaining the earlier improvements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.