When Should AI Evaluation Measure More?
Abstract
When an evaluation result is uncertain, the usual remedy is more of the same evaluation. We argue that evaluation is instead a target-conditioned design problem: an evaluator can change what is measured, how the responses are analyzed, or which target and reporting rule the evidence is meant to support, and these levers should be compared at matched cost. Across vision and language-model response banks, we find that the useful action depends on the bottleneck. In one matched-cost vision comparison, the same added budget reduced target error from 1.10 to 1.05 percentage points when spent on deeper replication of existing conditions, but to 0.56 when spent on target-relevant joint conditions. More broadly, reanalysis can outperform substantially more calls under a plug-in analysis for lower-tail targets but not for mean targets, and a rule accurate for a scalar summary can fail on the condition-level responses behind it. In a held-out check, a comparison frozen on development models selected released estimators that reduced equal-query error on held-out models when a historical calibration bank was available. We formalize the distinction between measurement support, statistical inference and decision targets, and provide Reassess, a toolkit for comparing these alternatives from response-level banks. The practical question is therefore not only how much more to evaluate, but what the next unit of evaluation should change.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.