acceptodds
Under review as a conference paper at ICLR 2027

VETO: Reward-Supervised World Models as Action Verifiers for Sample-Efficient RL

Abstract

Fine-tuning pre-trained manipulation policies with reinforcement learning (RL) is a promising route to real-world adaptation, but its sample efficiency remains the limiting factor on real robots. A key source of wasted interactions comes from credit-assignment lag, i.e., RL policies repeatedly reattempting failure modes they have already observed because observed outcomes must be propagated through a slowly converging temporal-difference (TD) critic, and then through the policy, before behavior changes. Model-based RL methods are often used to improve sample efficiency since they rely on supervised learning, but they can introduce model-error bias into policy gradients. Our key insight is that action ranking need not rely on TD learning: a latent world model trained with *supervised reward prediction* can provide useful rankings with fewer transitions than TD-trained -functions. We introduce Verifying Exploratory Trajectories Online (), which uses a continually learning, lightweight world model as an online action verifier. The policy proposes candidate action chunks, and the verifier scores them by rolling out latent dynamics and predicting rewards to veto poor action candidates. Because policy and critic targets use real transitions rather than imagined rollouts, world-model errors do not enter the Bellman backup, allowing for asymptotic optimality. Across 7 tasks in 2 simulation suites and on real-world YAM and Franka platforms, adapts in fewer environment interactions than off-policy and model-based baselines and improves pre-trained policies in the real world.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.