Scoring the Same Outputs: How Answer-Extractor Choice and Consistency Change Benchmark Comparisons
Abstract
Free-response benchmarks do not score model text directly: an answer extractor first decides which number counts as the answer, so two pipelines scoring the same outputs can report different gaps. We make that scoring step the object of study. For a frozen class of seven deterministic extraction rules, we compute every gap fixed outputs can yield when one shared rule scores every response, and how far the range expands when the rule may vary per question or per system—oracle stress tests of assignment consistency. A registered reanalysis of a disclosed development cohort (22 comparisons, four 1–2B families) and a preregistered held-out confirmation locate both robustness and fragility, and the pipeline falsifies its own development sign flip. On 193 harder ContextMATH pairs the audit separates two regimes: Qwen2.5's scenario effect has no unique value until the extraction rule is declared (a shared-rule range of −13.0 to +12.4 points at 7B, −14.5 to +8.8 at 32B—rule dependence, not a confidence interval), while four newer checkpoints from two families earn exact within-class robustness certificates: one identical negative gap under every rule and assignment. Bigger models alone do not remove scoring ambiguity. Declare one extractor before scoring, apply it to every system, and report the shared-rule range over a prespecified class.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.