acceptodds
Under review as a conference paper at ICLR 2027

Strict Answer Grading can Reverse its Error during Fine-turing

Abstract

Reasoning models are trained, selected and compared by the accuracy of the last boxed answer they write, the default in widely used frameworks. We show that this rule can misreport progress by tens of points, and in opposite directions within one fine-tuning run: early on it credits boxes in reasoning that never finished (at temperature 1 it reports 24.1% on AIME 2025 when only 0.9% of responses deliver a correct answer), and later it misses correct answers stated without a box. Scoring instead the answer each finished response delivers (delivered accuracy) changes four experimental conclusions: a pre-registered fix that raised strict AIME accuracy after an empty thought by 31.0 percentage points did not improve delivered accuracy. On GPQA, strict grading also puts four of twelve public reasoning models at 0.0–11.2% while their delivered accuracy is 35.6–63.6%, and reorders 13 of 66 model pairs. An exact identity, valid for any model, writes the gap as credit for boxes in unfinished reasoning minus correct answers left unboxed, so two counts that need no judge bound it. The bound holds in 237 of 238 cells across three students and twelve public models, the larger count gives the error's sign in 99 of 110 cells where both are nonzero, and removing only the boxes from otherwise identical targets reproduces the late error in a second model (two seeds). Below a combined 2 points, the counts certify a strict score to within about 2 points without a judge; above it, we give a reporting protocol.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.