acceptodds
Under review as a conference paper at ICLR 2027

Learning to Scale at Test Time: RL from LLM Feedback Loops

Abstract

Test-time scaling improves large language model reasoning by allowing models to explore, revise, and aggregate multiple attempts. Reinforcement learning (RL) offers a way to develop these capabilities, but training on extensive interaction trajectories can be costly, and the value of different training data remains poorly understood. In this work, we investigate how to select useful RL contexts and learn from offline solution attempts and feedback. We construct a dataset of 200K offline trajectories by iteratively generating solutions and feedback for 6K mathematics problems. Through extensive post-training experiments on this dataset, we identify important factors such as solution diversity, loop count, and surprisingly irrelevant factors such as orchestration topology. To put these findings into practice, we select trajectories from the latest revision round that satisfy accuracy and diversity constraints for each problem. Using these selected data, we train Qwen3.5-4B without additional supervised fine-tuning, achieving 82.5% accuracy on IMO-AnswerBench with test-time scaling, surpassing DeepSeek-Math-V2 (75.8%) and Gemini Deep Think (IMO Gold, 80.0%). Our trained Qwen3-4B model also matches QED-Nano's accuracy with approximately half the inference compute.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.