Changing the Candidates, Not the Vote: A Repair Portfolio for Self-Consistency under a Per-Call Token Cap
Abstract
Self-consistency (SC) votes over sampled reasoning chains, but under a hard per-call token cap many chains never finish, and a system that discards unfinished generations has nothing to vote for. We study an alternative at an equal number of calls: replace three of eight SC samples by a greedy chain that does not vote and two repairs of it, and return the SC plurality when any SC sample answers and the repair plurality otherwise. With Qwen3.5 models on MATH Level 4–5, this frozen configuration raises accuracy over SC@8 by 10.0 percentage points with 95% confidence interval in an exploratory study at a 4,096-token cap, by 3.7 in a pre-registered evaluation of a 475-question archive, and by 9.4 and 11.7 in a pre-registered replication on 100 unused questions at caps of 4,096 and 16,384 tokens. The gain comes from answer supply: it arises where no SC chain finishes, because the repairs finish more often. It persists when answers are recovered from unfinished chains and when SC is sampled at the repairs' temperature or with a budget instruction in every prompt, and the 15-call portfolio exceeds SC@16 at the same cap in the original study and at both prospective caps.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.