acceptodds
Under review as a conference paper at ICLR 2027

Evaluator validity is a property of the (evaluator, task) pair

Abstract

Cheap stand-ins for expensive robot policy evaluation, whether simulators, learned world models, or vision-language judges, rest on one load-bearing argument: a high rank correlation between the cheap and the expensive ordering on a handful of policies. Careful papers check more; none of those checks replaces it. We show this evidence does not support the conclusion drawn from it, and say what to report instead. Using a deterministic ground-truth evaluator, so that every disagreement is judge error and nothing else, we score 12 (judge, task suite) pairs built from three vision-language judges and four LIBERO suites, holding the prompt, the policy family and the construction of the policy set fixed. Validity turns out to belong to the pair: within a single suite the probability of choosing the truly best policy varies by 0.835 depending only on the judge, and no judge is best on every suite. Rank correlation tracks none of it. It exceeds 0.82 in 10 of the 12 cells while the probability of picking the best policy inside that group runs from 0.212 to 1.000, and its sign changes with how the policy set was assembled and with the wording of the prompt. In the practitioner’s own units, following one of these certified evaluators gives up as much as 14.7 points of true success rate, and the worst evaluator we measure is beaten by a coin. The same failure appears between two independently released checkpoints, where one judge is significantly backwards on a true gap of 12.6 points. Our coverage is narrow, three judges and two public checkpoints, so we map this failure rather than bound it. We give the perfect-evaluator ceiling that makes any reported correlation readable; on a published leaderboard a flawless evaluator would score only ρ = 0.500 on LIBERO-Long at the budget the field uses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.