SELFOCUS: Mitigating Systematic Selection Bias in Inference-Time Scaling
Abstract
Inference-time reasoning methods often rank sampled chains of thought (CoTs) using within-trace token probabilities. Final-answer frequency, however, is a many-to-one trajectory marginal. Under an idealized outcome-only, KL-regularized trajectory model, we prove that self-consistency need not recover the highest-reward answer because its ranking also depends on the total reference-policy mass of all trajectories leading to answer . We identify a sufficient high-dimensional regime in which heterogeneous erroneous paths make a wrong answer more frequent than the uniquely rewarded answer, and demonstrate this distributional shift in a controlled multi-step experiment. We then introduce SELFOCUS, a training-free aggregator that uses cross-CoT semantic recurrence to estimate each reasoning path's deviation from the trajectory population and partially mitigate the influence of on inference-time answer selection. It clusters reasoning steps, accumulates question-local log-frequency offsets along each path, and selects the answer minimizing the ratio of mean cumulative offset to answer frequency. Across four model families and seven benchmarks, SELFOCUS improves paired majority vote in 15 of 17 model–dataset settings without token log probabilities, an answer-correctness verifier, or additional training. On larger models, measured compute also shows substantially lower GPU-hour usage than token-log-probability-based selectors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.