acceptodds
Under review as a conference paper at ICLR 2027

Accepted but Invalid: When Reasoning Benchmarks Credit Wrong Answers

Abstract

When a reasoning benchmark marks an answer as correct, we usually assume that the submitted answer actually satisfies the task. We show that this assumption can break, even while the evaluator runs exactly as implemented. In MathConstruct, for instance, an archived model answer repeats an index in a required permutation, yet the released checker accepts it and awards credit. To see how far the problem extends, we craft small, explicit violations of the written task requirements and pass them through the released evaluation code. Across seven MathConstruct tasks, the released checkers accept all 127 applicable invalid constructions in our fixed cross-instance test, and the erroneous acceptance of five invalid archived model answers has already entered the task-weighted benchmark scores. DCP-Bench-Open reveals a further and subtler failure: omitted answer fields are left free for the evaluator’s solver to complete. In the tested release, the evaluator accepts an empty JSON object for all 164 problems in our controlled completeness test, and accepts 735/735 incomplete outputs overall. Together, these findings show two ways an evaluator can award credit for a task the model never completed, i.e., by checking too little of the submitted answer, or by completing part of the answer itself.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.