acceptodds
Under review as a conference paper at ICLR 2027

The Underdetermined Stream: What a Single-Gold Score Actually Measures on Contested Material

Abstract

Benchmarks for mathematical reasoning compare a model's answer with a single gold string, and on material where several forms of the same answer are equally correct that comparison is measuring something other than capability. The field treats the mismatch as a data-quality bug for a better normaliser to repair, which presupposes that one of the two golds is wrong; when both are correct there is nothing to fix, and what the score measures instead has never been derived. The obvious remedy — publish a correction factor beside the score — presupposes that the multiplier is a property of the corpus, and latent-variable estimators of a hidden true label cannot price their own error. We derive the identity: an exact-match score on contested material is capability multiplied by the rate at which the model writes the one form this corpus shipped as gold, a product whose second factor belongs to the model-corpus pair rather than to either alone. Capability is nonetheless exactly recoverable without retraining, by scoring the same generations against each convention's gold in turn and summing, under two premises that are measured rather than assumed. On four arms of one corpus differing only in the order training data reached the optimiser, capability holds inside the 0.0270 measurement floor while the allocation moves 0.4621 — the policy at F=612.18 against the capability sum's 1.63 — so a single-gold benchmark reports the same capability 4.75x apart depending only on how the model was taught to write. The estimator returns an estimated capability of 0.4452 against a directly measured 0.4522, and that 0.0070 gap is exactly the support leak the identity predicts (0.00705), so its error is always downward and always priced by a quantity the same run reports. The correction-factor shortcut is registered, run and rejected: the model's policy correlates with per-item solvability at r=-0.257 on 24 of 24 arms, so the multiplier is a property of the model as well, and the effective enumeration it hides runs 1.149 to 1.811. The rule is to publish the summed score, which costs extra scoring passes over generations already produced and no new training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.