Benchmark Graders Miss What Agents Leave Behind
Abstract
Benchmark graders can award passing scores to outputs that violate task requirements. Repairing such errors can change which agent is recommended, but the value of that change depends on performance on new outputs. We study this distinction using repeated coding-agent executions, archived public code generations, and a documented banking-benchmark correction. Additional checks reveal sensitive keys retained in nested outputs and invoices reset after an intervening payment. We compare the original-score and corrected-score choices on the same corrected target in another execution. Across four new model environments, mean selection benefits include gains, no change, and losses. The negative mean persists in one environment when the disputed migration check is omitted. On public code generations, selectors frozen before held-out scoring reduce mean pairwise loss from 0.098 to 0.052 passed tasks, while the full-pool leader remains unchanged. The banking correction also reduces mean selection loss while retaining adverse transfers. These results separate checking task requirements from evaluating the consequence of a revised recommendation. We release deterministic checks, paired score records, and a calculator for reproducing these comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.