acceptodds
Under review as a conference paper at ICLR 2027

Correlation Does Not Certify a Policy Choice: What Benchmarks Should Publish Instead

Abstract

A world model that correlates at .97 with real robot trials can still pick a worse policy per task than the best single policy on that board. Evaluators like it are validated against that printed table of real outcomes by publishing a correlation; we compute what one certifies about the choice. That certificate is an attained maximum over a cone of score matrices, depending on the decision. For ranking the policies and shipping one, the usual claim, four of eleven stay within the cost of shipping the best one, and shortlisting two for hardware improves eight pairs. For choosing a policy per task, a decision we price against a correlation rather than observe, exactly one is adequate; ten leave open a regret above that baseline and two certify nothing at all. A slightly higher correlation would not rescue the finer decision. Two numbers per task and two for the board bracket the correlation ruling out a five-point regret, wherever the ranges clear the budget. That bar lies between .978 and .998 here and above .97 on 498 of 499 random boards, and its mean did not fall as more tasks of the same kind were added. A different summary converts better, and two of these benchmarks already publish it. An evaluator whose margin-weighted rank violation is over policies gives up at most choosing a policy per task, and half that where its taskwise argmax is unique. Priced exactly, the statistic beats the correlation on all eleven pairs, but for a board-plus-noise evaluator it does nearly as well: the finding is the conversion, not the corpus. With no board, the correlation's bound gives nothing on six settings of thirteen, the statistic's on one. We recommend publishing it, with trial counts, per-task margins and ranges, and a declared accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.