Controlling Length Growth in Reinforcement Learning for Code with Adaptive Budget
Abstract
Reinforcement learning (RL) improves the reasoning accuracy of large language models but also inflates response length. We study how to curb this growth without suppressing useful reasoning, in competitive programming with training contexts of up to 256k tokens. We find that length grows most, proportionally, on problems the initial model already solves reliably, while accuracy gains concentrate on hard problems. Length penalties are a common remedy, but their shape and strength do not map directly to response length, which makes them hard to tune. Explicit rollout budgets instead limit length in tokens, yet penalizing truncated responses under a fixed budget shortens outputs at the expense of hard problems. Since easy and hard problems call for very different lengths, no single budget serves both. We therefore propose adaptive budgets, which, unlike length penalties, express each prompt's limit in tokens. Each prompt is capped at a multiple of the initial model's median successful length, rarely solved prompts keep the full budget, and responses truncated at the cap are masked from the loss rather than penalized. Compared with a fixed-budget baseline, adaptive budgets cut mean response length by over 30% on LiveCodeBench with no loss in accuracy, and by 25% on the harder LiveCodeBench-Pro at a 1.5-point accuracy cost, improving on the accuracy–length tradeoff of tuned length penalties. Savings are largest on initially solved problems and transfer to AIME 2025/2026, with about 15% fewer tokens at similar accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.