Learning to Scale at Test Time: RL from LLM Feedback Loops
Abstract
Test-time scaling improves large language model reasoning by allowing models to explore, revise, and aggregate multiple attempts. Reinforcement learning (RL) offers a way to develop these capabilities, but training on extensive interaction trajectories can be costly, and the value of different training data remains poorly understood. In this work, we investigate how to select useful RL contexts and learn from offline solution attempts and feedback. We construct a dataset of 200K offline trajectories by iteratively generating solutions and feedback for 6K mathematics problems. Through extensive post-training experiments on this dataset, we identify important factors such as solution diversity, loop count, and surprisingly irrelevant factors such as orchestration topology. To put these findings into practice, we select trajectories from the latest revision round that satisfy accuracy and diversity constraints for each problem. Using these selected data, we train Qwen3.5-4B without additional supervised fine-tuning, achieving 82.5% accuracy on IMO-AnswerBench with test-time scaling, surpassing DeepSeek-Math-V2 (75.8%) and Gemini Deep Think (IMO Gold, 80.0%). Our trained Qwen3-4B model also matches QED-Nano's accuracy with approximately half the inference compute.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.