acceptodds
Under review as a conference paper at ICLR 2027

Beyond Plan Selection: Attributing and Improving Execution in RLVR

Abstract

Does reinforcement learning with verifiable rewards (RLVR) improve reasoning by selecting better plans or by executing them more reliably? We study this question through a matched plan-conditioned interface that allows Base and RLVR-trained selectors and executors to be interchanged while keeping the execution context fixed for each task–plan pair. Using a 1.5B model in two controlled two-plan environments, we separate realized performance contributions from available routing opportunity and the effects of targeted training. In arithmetic, the prespecified component-replacement analysis assigns 2.27 percentage points more improvement to execution than selection (95% CI: 1.14–3.41), although a common-repeat sensitivity analysis does not establish this ordering. Motivated by this diagnosis, we evaluate Execution-Focused RLVR (EF-RLVR), which removes the direct selector-policy loss while retaining shared parameters and KL regularization.EF-RLVR improves accuracy over Standard Joint RLVR by 1.28 points across eight fresh paired training seeds (95% CI: 0.77–1.79). A factorial follow-up and a supplementary direct comparison support a larger benefit from selector-term removal than executor up-weighting. In Search, input-dependent routing remains beneficial after training under both structured and final-answer scoring, but neither the original selector objective nor a uniform-action-weighting variant learns the conditional routing rule. These results show how component attribution can guide a useful training intervention, while demonstrating that available selection opportunity need not translate into learned routing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.