acceptodds
Under review as a conference paper at ICLR 2027

Reliability-Aware World Model for Robot Policy Improvement

Abstract

Robot policies can improve by practicing inside a learned world model, but the model is wrong in places, and practicing there teaches the policy behaviors that fail on the real robot. Fixing the model needs new real data, and the usual way to get it is to run the policy on the robot and refit; most of those rollouts, however, land on interactions the model already predicts well, and few reach the places where it is wrong. We present a method in which the world model's own prediction says how far it can be trusted. A reliability score, computed from the self-consistency of the flow model's velocity field and from whether the predicted motion follows the commanded action, is used twice: it weights each imagined policy update, and it selects the interactions at which real data is collected. The world model is a video transformer with a dynamics head and an action- and contact-conditioned value head on a shared backbone, and its action-value gradients drive policy improvement. A flagged interaction is reached on the robot by restoring the logged start and replaying the action prefix; the real outcome refits the model before the next round. On three simulation and three real-robot tasks, two rounds of 20 real rollouts improve success over the initial policy by 15.8–17.1 and 25.0–28.1 points and are ahead of DSRL and a VLAW reproduction on every task at the same rollout budget. Controlled studies show that both uses of the score contribute, and that the collected data corrects the model where it was flagged.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.