When Grading Is Not Mode-Invariant: Answer-Extraction and Normalization Errors in Think/No-Think Comparisons
Abstract
Choosing whether to use an explicit thinking phase requires a reliable estimate of its reported (grader-measured) accuracy gain. Yet on identical MMLU outputs, a vendored Qwen multiple-choice extractor applied to full transcripts reports a think-minus-no-think gap of -15.7 points, while two other mainstream extractors report +17.0 to +17.3; the disagreement persists on the complete MMLU test set. Scoring only the answer segment flips the MMLU gap's sign; one GSM8K extractor's failure persists under the same change. We audit 200,328 Qwen3-8B candidates from five benchmark-derived item sets, using 595 stratified reference labels (LLM-prescreened, human-checked) and separate follow-up and transfer studies. Source-code interventions identify repeated answer-trigger phrases that, in affected outputs, truncate the transcript before final-answer extraction; on 240 Qwen candidates of a separately sampled MMLU follow-up, we evaluate the answer-segment rule whose population verdicts were recorded before that draw. Against the returned question-answer key, the repair removes 47 reference-grading errors and introduces four among the 51 sampled think verdicts that change (51 and zero against benchmark gold). The estimated population error reduction for the 2,560 think outputs has a conditional 95% interval [20.1, 43.9] points; no-think verdicts are unchanged. This does not establish the semantic accuracy ranking. On GSM8K, a last-number extractor's reversal survives the same repair. An exact error decomposition relates mode-specific errors to the reported gap; conditional bounds specify when the evidence supports a comparison; agreement alone cannot certify validity. These results support checking repairs against reference labels and reporting the remaining uncertainty; inference concerns fixed corpora and reference conventions, without adjustment for historical search, undetected label errors, or the post-discovery assembly of the mainstream panel.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.