acceptodds
Under review as a conference paper at ICLR 2027

Benchmark Graders Miss What Agents Leave Behind

Abstract

Benchmark graders can award passing scores to outputs that violate task requirements. Repairing such errors can change which agent is recommended, but the value of that change depends on performance on new outputs. We study this distinction using repeated coding-agent executions, archived public code generations, and a documented banking-benchmark correction. Additional checks reveal sensitive keys retained in nested outputs and invoices reset after an intervening payment. We compare the original-score and corrected-score choices on the same corrected target in another execution. Across four new model environments, mean selection benefits include gains, no change, and losses. The negative mean persists in one environment when the disputed migration check is omitted. On public code generations, selectors frozen before held-out scoring reduce mean pairwise loss from 0.098 to 0.052 passed tasks, while the full-pool leader remains unchanged. The banking correction also reduces mean selection loss while retaining adverse transfers. These results separate checking task requirements from evaluating the consequence of a revised recommendation. We release deterministic checks, paired score records, and a calculator for reproducing these comparisons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.