SCALABLE: Scaling CApabilities Beyond the Best-of- Limit with Adaptive Boundary LEarning
Abstract
Reinforcement learning with verifiable rewards (RLVR) often improves the single-sample reasoning performance of large language models (LLMs), yet whether the long-run capability growth can surpass the frontier reachable by the initial policy through best-of- sampling remains debated. In this paper, we investigate the conditions under which RLVR can sustain capability growth beyond this apparent ceiling. We study problems with a controllable difficulty scale and fine-grained, verifiable solution-quality evaluation, with a focus on combinatorial optimization (CO). We observe a sharp performance decline as task difficulty increases, which we characterize as an empirical capability boundary. Under RL with uniform sampling, this boundary shifts toward harder instances, and its movement can be tracked using intrinsic or extrinsic signals. Based on these observations, we propose Adaptive Boundary Reinforcement Learning (ABRL), a novel learning strategy that dynamically concentrates training near the model's moving capability boundary. Across three CO tasks, ABRL sustains capability improvement, while experiments in other reasoning domains provide supporting evidence of its effectiveness beyond CO. At the largest problem sizes in our best-of- evaluation across the CO tasks, ABRL achieves average normalized best-of-1 performance of 0.5548, approximately the initial policy's best-of-4096 performance. In addition, we find that the number of training steps required to reach target performance exhibits an empirical power-law relationship with problem size across a broad range of problem settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.