Accuracy Is Not Stability: Answer Turnover in Test-Time Scaling
Abstract
More inference compute can repair errors but also discard answers that were already correct. We study these opposing flows, repair and corruption, under nested best-of- selection. Two populations can share identical accuracy and oracle-coverage curves at every width, yet lose correct answers at different rates. Exact score-record laws explain this gap: if correctness at extreme verifier scores tends to , each flow from to candidates is asymptotic to as grows. Adjacent turnover therefore vanishes, whereas multiplicative budget increases retain positive turnover even when accuracy rises monotonically. We derive unbiased finite-pool estimators of both flows with a Rao–Blackwell variance guarantee, and formulate answer preservation as a whole-policy risk constraint. Across five model–task–verifier settings, expanding from one to candidates improves accuracy, yet loses a previously correct answer in – of evaluations. At a mean budget of about candidates, a preservation policy attains similar observed accuracy to randomized widths while reducing conditional corruption by percentage points; an adaptive comparator attains higher accuracy with more corruption. These results separate accuracy improvement from answer preservation and quantify their trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.