The Sampling Ceiling: Why Test-Time Self-Consistency Saturates and an Oracle Rule for Allocating Samples
Abstract
Self-consistency can sample a correct answer and still vote for a wrong one. We study how confidence in the current answer differs from the value of drawing more samples. In modular addition, 40 samples contain a correct answer on 95.34% of instances, yet voting accuracy converges to 63.32% (three-seed means). In our systematic-stochastic mixture, voting can remove random errors on some instances, while on the rest the model always repeats one wrong answer. Each extra pair of samples then adds less majority-vote accuracy than the last, so a greedy oracle allocation is optimal for equally weighted queries with known per-query mixture parameters. When these are unknown, Correlation-Aware Self-Consistency (CASC) uses labeled population data to score the current answer's correctness and the risk of a systematic error. On public Grade School Math 8K (GSM8K) answer sets, the resulting tradeoffs depend on the model: default CASC uses 3.67 answers versus Early-stopping Self-Consistency's 8.24 on GPT-4, with a paired 95% interval for the accuracy difference that includes zero. Removing CASC's systematic-error term leaves its GPT-4 predictions and stopping times unchanged; Adaptive-Consistency's Beta rule is more accurate and cheaper on GPT-3.5. The oracle ranks queries by the gain from additional samples, and the stopping comparisons show that scores for the current answer do not measure that gain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.