Accepted but Invalid: When Reasoning Benchmarks Credit Wrong Answers
Abstract
When a reasoning benchmark marks an answer as correct, we usually assume that the submitted answer actually satisfies the task. We show that this assumption can break, even while the evaluator runs exactly as implemented. In MathConstruct, for instance, an archived model answer repeats an index in a required permutation, yet the released checker accepts it and awards credit. To see how far the problem extends, we craft small, explicit violations of the written task requirements and pass them through the released evaluation code. Across seven MathConstruct tasks, the released checkers accept all 127 applicable invalid constructions in our fixed cross-instance test, and the erroneous acceptance of five invalid archived model answers has already entered the task-weighted benchmark scores. DCP-Bench-Open reveals a further and subtler failure: omitted answer fields are left free for the evaluator’s solver to complete. In the tested release, the evaluator accepts an empty JSON object for all 164 problems in our controlled completeness test, and accepts 735/735 incomplete outputs overall. Together, these findings show two ways an evaluator can award credit for a task the model never completed, i.e., by checking too little of the submitted answer, or by completing part of the answer itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.