Where Best-of-\(N\) Gains Are Lost: The Limits of Self-Sourced Code Selection
Abstract
Sampling multiple candidate programs is a standard way to scale inference-time performance for code generation. Although increasing the sample count makes it more likely that a correct program is present, realizing this gain requires reliably identifying a correct one. We study this problem when selection relies on signals available from the generator itself, including candidate agreement, self-generated tests, and self-judgment. We find a striking mismatch between where resampling creates opportunity and where these signals are effective: across four models on LiveCodeBench, 78–86% of the gap between single-sample and Best-of- performance is concentrated on minority-correct problems, where only 1–5 of 10 candidates are correct. This is precisely the regime in which self-sourced selectors are weakest, and none significantly outperforms a simple filter that executes the worked examples provided in the problem statement. We isolate two mechanisms behind this shortfall. First, even an oracle behavioral partition leaves plurality selection weak: identifying behavioral classes does not identify the correct class when correct behavior is a minority. Second, generated tests are often discriminating but mislabeled; validating their expected outputs raises a simple test filter 5.7–11.9 points above the example filter on minority-correct problems. Together, these results show that producing additional distinctions among candidates is not enough: the central challenge is grounding those distinctions in correctness. Effective inference-time scaling through resampling therefore requires selection evidence whose connection to correctness can be independently grounded.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.