acceptodds
Under review as a conference paper at ICLR 2027

Generative Models as Online Policy Correctors

Abstract

How can an agent learn efficiently online? Reward-based reinforcement learning can require costly exploration to discover better behaviors in open-world environments with sparse rewards. We propose an alternative in which an agent learns not only by reinforcing successful behavior but also by correcting mistakes in imagination online. We introduce \method, a learning algorithm that fully closes the loop in online corrective learning using pretrained generative models as online correctors. Given a policy rollout, \method identifies failures and generates counterfactual rollout continuations that correct them. The policy is then finetuned on these corrected rollouts and redeployed in the environment to collect new experience, repeating this cycle online. On Minecraft ObtainDiamond, \method substantially improves both behavior-cloned and vision-language-action policies. By applying pretrained priors at the learner's own failures, ReVision achieves returns comparable to PPO using 5 times fewer online environment steps (6M versus 30M). Controlled comparisons show that repeatedly correcting the updated learner yields greater improvements than distilling teacher demonstrations or correcting a frozen learner. Further experiments on manipulation tasks like RLBench also demonstrate improvements in pretrained manipulation policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.