Curse of Corpus: Do Models Reason Beyond Counting And Composition?
Abstract
Process reward models are often evaluated by pooling all labeled steps into one AUROC. We study this practice under the absorbing-label convention, which labels the first incorrect step and every later step as incorrect. Pooled AUROC then mixes two nuisances: comparisons across heterogeneous sources and an association between position and incorrectness. Both are measurable exactly, because a rank statistic is a sum over positive-negative pairs: on our held-out set only 0.59% of the 371,765,485 pairs compare steps of one source at one index, and 93.84% cross sources. A scorer that reads no text and knows only the corpus therefore reaches pooled absorbing AUROC 0.9140, against 0.9151 for a fine-tuned 31B verifier. Across five trained verifiers, held-out raw pooled AUROC exceeds mean-macro AUROC by 0.0963-0.1244. Across seven corpora, a text-blind relative-position baseline obtains absorbing-label pooled AUROC from 0.6370 to 0.9768; imposing absorption raises it by as much as 0.3124 without changing text or scores. We evaluate first-error detection with , which excludes post-error suffixes and compares errors with correct steps at the same index. Absolute index alone therefore receives exactly 0.5, mean-macro and weighted-macro variants also exclude cross-source comparisons, and adding length strata returns relative position to 0.5. The estimands then disagree about direction and not only magnitude: averaged over the first and last five checkpoints of a matched 2,300-step run, every ranking estimand rises with an interval excluding zero, pooled AUROC by 0.0708 and weighted-macro by 0.0363, while accuracy at naming the step that first broke falls by 0.0335, interval . In a forced-verdict diagnostic, pooled raw AUROC places Qwen 0.0847 above Gemma whereas source-aware first-error intervals include zero; conversely weighted-macro favors Muse over GLM by 0.0597 with interval while pooled raw AUROC is inconclusive.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.