acceptodds
Under review as a conference paper at ICLR 2027

How Far Can Reasoning Alone Take You? Graph-Based Candidate Selection for Test-Time Scaling on Code

Abstract

Large reasoning models spend most of their test-time compute on long chains of thought, and on competitive programming the quality of that reasoning — rather than implementation ability — is the dominant bottleneck. Given sampled candidates per problem, the question is which one to trust. We ask how much of that selection problem is solvable from the reasoning trace alone: no code execution, no answer consensus, no token probabilities. We condense each trace into a summary, lift the summary into a reasoning graph whose topology encodes steps, dependencies and case analysis, and train a graph neural network (GNN) to predict correctness from the graph. On LiveCodeBench-v6, under a leakage-free 10-fold cross-test protocol in which every problem is scored only by models that never saw it during training or checkpoint selection, per-fold ensembles recover 0.29 of the gap (+9.5 points absolute over random selection), against 0.12 for a frontier judge reading the same traces and 0.16 for a topology-free MLP over the same summary embeddings; single models recover 0.19, so seed-ensembling — effective here because a small training set and aggressive dropout yield weakly correlated models — accounts for a third of the effect, and the worst of the ten folds still recovers 0.16. A comparison of seven checkpoint-selection rules shows that selecting on a smoothed validation curve dominates raw peak selection in both mean and worst-case usefulness. As a transfer example, on a harder LiveCodeBench-Pro subset the classifier–selector cascade recovers . We then characterize the ceiling of reasoning-only selection: 94.5% of incorrect candidates textually claim success, and most residual errors are near-ties where candidates share an algorithm and differ only in implementation details absent from the trace — so much so that the same judge, held fixed, recovers three times as much of the gap from the candidates' code as from their reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.