Right Proof, Wrong Theorem: A Relational Bottleneck in Verifier-Guided Scaling
Abstract
Best-of-N selection samples N candidate solutions and returns the one preferred by a verifier. It improves only if the selected candidate solves the requested problem. A coherent proof or program may instead solve a nearby problem, a failure we call a relational mismatch. We certify such errors in code and mathematics outputs using claim neighborhoods frozen before artifact inspection. Across two frozen generators, the upper score tail of our primary verifier contains fewer local defects but more relational mismatches. For the primary generator, raw best-of-N validity first rises and then falls: from 88.41% at N=32 to 80.34% at N=256 in code, and from 82.03% at N=16 to 76.56% at N=64 in mathematics. Exact finite-pool selection weights forecast these two curves with weighted mean absolute errors of 0.71 and 1.09 percentage points. A development-frozen interior-rank selector then improves fixed-budget validity by 4.49 to 8.98 percentage points across four domain-generator settings, without additional generation or verifier calls. Finally, in the primary code setting, relational training improves held-out claim-artifact binding and natural deployment; replacing ordinary relational ranking with a crossed relational objective adds 1.65 percentage points. These results show that recognizing a good artifact is insufficient: reliable verifier-guided scaling also requires binding the artifact to the claim it must satisfy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.