acceptodds
Under review as a conference paper at ICLR 2027

CANDIDATE ADMISSION GATES FOR GENERATIVE REASONING: DIAGNOSTICS AND VMGR PLANNING

Abstract

Which change does a performance gain in generative reasoning identify? We study this question through comparison diagnostics and a controlled generativeplanning method. Three formative comparisons of autoregressive and iterative systems expose distinct population, competence, measurement and provenance failures in code repair, GF(2) reasoning and minimal planning. We organize these failures into six candidate admission gates for bounded, verifiable tasks. A claim-specific state representation and CPU replayer distinguish blocked attribution from reportable system outcomes; the cases motivate, but do not independently validate, the procedure. A later planning diagnosis finds that selecting preferred demonstrations also changes candidate coverage. VMGR separates the two: verified trajectories define shared support, budget-conditioned completion costs define soft targets, and a uniform floor retains every candidate in a root-balanced loss. With one 0.5B language model and three paired rank-64 fits, VALUE exceeds same-support UNIFORM by 3.646 percentage points on 768 new exact task identities. Later search recovery reaches 99.94% on 600 smaller public tasks, but almost removes the weighting gap. On all 110 tasks in a harder public configuration, frozen VALUE and same-search UNIFORM reach 62.73% and 52.12%. The 10.606-point difference has a conditional task-cluster 95% interval of [4.848, 16.667], alongside 14.34% less inference token-work. Thus the strongest method evidence is a same-support supervision gain, not the larger gap against poorly transferring SFT controls. The study supports a scoped planning improvement and useful comparison distinctions, not universal gate validity or architecture-level superiority.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.