Same Outputs, Different Winners: Evaluator-Dependent Comparisons in Code Benchmarks
Abstract
AI coding research relies on benchmarks to compare code models and coding agents, assess progress, and select research baselines. Yet changing the evaluator can reverse statistically supported leads even when code outputs remain fixed: 60 of the 66 supported reversals we observed had point-estimate advantages of at least 2 percentage points in both opposing directions. To characterize this dependence, we hold tasks, system configurations, and generated outputs fixed and study evidence reinterpretation, criterion strengthening, and specification revision through evaluator contrasts in SWE-bench Verified, EvalPlus, and LiveCodeBench, respectively. Within each study, joint inference over all system pairs and evaluator conditions provides a common evidence standard, separating support lost through an expanded inference family from support lost after evaluator changes. We find that only 38.5%-55.0% of originally supported leads retain support for their original direction across evaluators; among cross-evaluator support losses, roughly two-thirds gain support for the opposite direction. The consequences also depend on where the affected relations lie: a leading group can retain supported leads over all remaining systems while its internal winner remains unresolved, whereas a single winner-critical conflict can block a unique-winner claim. We therefore propose Claim-first interpretation, which begins with a target claim and its evaluation scope, examines the pairwise relations required to support it, and determines what conclusion the available evidence supports. In practice, the framework guides researchers in assessing progress and selecting baselines and shortlists, authors in reporting conclusions and their evaluation conditions, and maintainers in reassessing critical comparisons after evaluator updates. It thus moves benchmark use beyond reading scores and ranks toward evidence-based scientific conclusions that can be reexamined as evaluation conditions change.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.