Inference-Time Amplification of Weak Reasoning Models
Abstract
How much of the capability of a reasoning model exposed by repeated sampling can an imperfect selector recover? We study this problem as \em inference-time amplification, separating \em coverage—whether a useful candidate is generated—from \em identifiability—whether it can be recognized among the alternatives. We show that better coverage alone does not guarantee better selection, while even imperfect local selection signals can be amplified through repetition. We also characterize proposal-side blind spots that additional sampling cannot overcome. On SWE-bench Verified, a single GPT-5.4 nano trajectory solves 67.0% of tasks, while eight proposals contain a correct patch on 79.0%; critic–comparator selection reaches 76.4%, recovering 78.3% of the oracle-exposed gain. The same pattern appears across proposer families and in mathematical reasoning, competitive programming, and formal proof. On MATH L4–5 and GSM-Plus, our selector recovers 85% and 77% of the oracle gap and substantially outperforms majority vote. Holding proposal coverage fixed, providing the selector with additional repository information also improves recovery. Conversely, when the available selection signal is weak, redundant, or misaligned, amplification can diminish or fail. These results identify selection quality as the key factor governing how much capability exposed by repeated sampling can actually be recovered.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.