Learn to Scale Reasoning at Test Time Via Reinforcement Learning
Abstract
Scaling test-time computation has emerged as a powerful paradigm for Large Language Models (LLMs). The effectiveness of sequential test-time scaling depends on whether models can effectively utilize additional inference budgets for further reasoning. However, we show that RLVR-trained reasoning models struggle to benefit from additional inference computation, even when larger budgets are available. In this work, we investigate whether RL training can teach LLMs to benefit from additional inference computation. To analyze how training changes the use of additional computation, we introduce state coverage as a proxy for in-context exploration and find that incentivizing additional computation during training can broaden coverage but also induce repetition. These findings motivate Learning to scale INference with lEss repetition (\method), a reward-shaping approach that encourages continued reasoning on unsuccessful attempts while controlling redundancy. Experiments show that LINE enables Qwen3-4B-Base to translate additional inference computation into accuracy gains, while also improving average mathematical reasoning accuracy by 4.4 percentage points over GSPO and strengthening out-of-domain generalization. Broader evaluations show that \method improves reasoning performance across multiple models and benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.