acceptodds
Under review as a conference paper at ICLR 2027

Correlation Does Not Certify Reliable Evaluation: Identifiability Limits of World-Model-Based Robot Policy Selection

Abstract

Policy evaluation with world-model rollouts and a VLM judge has been reported to correlate strongly with real-world success rates. We ask whether this correlation certifies sensitivity to the consequences of actions. Empirically, it does not. The evaluator's scores track simulator success rates, yet barely change under interventions that make the task fail in the simulator, and rise under a perturbation that lowers success rates. Most of the score variance across candidates is instead explained by similarity to the world model's training demonstrations. We then prove that, without additional real-world trials, no evaluation procedure can reliably identify the better of two candidates that differ only in an action unsupported by the demonstrations, even given a pretrained world model, the candidates' weights, and a perfect judge. Identifiability depends on , the expected number of demonstrations containing that action, rather than on , and perfect rank correlation can coexist with no response to such interventions. Finally, we propose ResGap, a coarse screening metric that needs neither world-model rollouts nor a VLM judge, ranks candidates more accurately than the world-model evaluator, and allows real-world trials to be concentrated on the top-ranked candidates. Correlation with success rates therefore does not certify reliable evaluation for world-model-based robot policy selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.