acceptodds
Under review as a conference paper at ICLR 2027

Learn to Scale Reasoning at Test Time Via Reinforcement Learning

Abstract

Scaling test-time computation has emerged as a powerful paradigm for Large Language Models (LLMs). The effectiveness of sequential test-time scaling depends on whether models can effectively utilize additional inference budgets for further reasoning. However, we show that RLVR-trained reasoning models struggle to benefit from additional inference computation, even when larger budgets are available. In this work, we investigate whether RL training can teach LLMs to benefit from additional inference computation. To analyze how training changes the use of additional computation, we introduce state coverage as a proxy for in-context exploration and find that incentivizing additional computation during training can broaden coverage but also induce repetition. These findings motivate Learning to scale INference with lEss repetition (\method), a reward-shaping approach that encourages continued reasoning on unsuccessful attempts while controlling redundancy. Experiments show that LINE enables Qwen3-4B-Base to translate additional inference computation into accuracy gains, while also improving average mathematical reasoning accuracy by 4.4 percentage points over GSPO and strengthening out-of-domain generalization. Broader evaluations show that \method improves reasoning performance across multiple models and benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.