acceptodds
Under review as a conference paper at ICLR 2027

Correlation Is Not Selection: Validating Automatic Evaluators for Robot Policies

Abstract

Testing robot policies on real hardware is slow, so automatic evaluators, such as reward models, vision-language judges, and world models, are increasingly used to choose which policy to deploy. What matters is the evaluator's top pick, yet evaluators are validated by how well their scores rank-correlate with success measured on real robots, over three to eight policies. We investigate whether such evaluators pick the right policy from recorded robot trials. We first compile an extensive set of 82 published evaluator–reality correlations from nine papers. Most of these rest on too few policies to support a general ranking claim: of the 36 computed over four or more policies, 25 have a 95% interval that includes zero. We then audit eight open-weight evaluator configurations on the same 5106 recorded episodes of seven robot-arm policies, against one set of double-blind human comparisons from the RoboArena benchmark. For each evaluator, we measure which policy it picks, how much worse that pick is than the humans' best, and how stable the pick is when the data is resampled. Three conclusions follow from these exploratory analyses. First, rank correlation does not tell which policy an evaluator will pick. Evaluators with the same correlation pick different policies, and they pick the humans' best policy in 16% to 80% of resamples, although these differences are not statistically established. Second, when evaluators miss the humans' best policy, they usually pick one almost as good, which the best policy would beat only slightly more than half the time. The exception, a score of how well a video's frames match the text instruction, picks the worst policy. Third, the best policy itself depends on whose preferences count and which policies compete. Counting each lab or human rater equally changes it, and adding one stronger policy drops the best evaluator's agreement to 20%. Evaluators should therefore be validated by the decision they support, against a named population and candidate set, not by correlation alone. We release the audit code, the frozen human ranking, and a report-card script that applies this audit to any new evaluator.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.