Answer Matchers Determine What Test-Time Sampling Appears to Do
Abstract
Test-time scaling is often framed as a policy problem. The system chooses how many samples to draw, when to stop, and whether to rerank, assuming that the measured outcome reflects the policy. On open-ended factual tasks, however, the measured outcome often reflects the answer matcher instead. We use the exact decomposition to split plurality error into non-coverage and aggregation failure. A matcher false negative enters both terms with opposite -dependence, creating an apparent conflict between candidate acquisition and selection. On Llama-3.1-8B/TriviaQA, one normalization step affects of all generations. Correcting it moves single-sample accuracy , turns total error from an increase () into a decrease (), and reverses the measured gain of adaptive stopping from to points. The distortion is predictable rather than idiosyncratic. For the exposed class of matching rules, it equals the fraction of a model's correct answers in the affected style times its accuracy. This relation has unit slope over 11 cells, replicates on EVOUNA's held-out human labels, and separates contemporaneous 8B open-weight models by an order of magnitude. Plain exact match is exposed in of cells and SQuAD-family normalization in none. We release the canary that detects the failure, which standard parsing metrics miss, together with two further cases in which a cheap scalar accounts for most of an effect attributed to an expensive mechanism. We compare process reward models with trace length and learned trajectory observers with prefix likelihood.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.