Continuous Test-Time Training for Large Reasoning Models
Abstract
Scaling test-time computation improves the reasoning performance of large language models (LLMs) without improving the models themselves. Test-time training (TTT) instead updates model parameters on unlabeled test problems. Without ground-truth answers, existing methods often use agreement or confidence as proxy rewards, achieving early gains that can plateau or reverse with further training. We investigate whether TTT can sustain improvement beyond the plateau of reinforcement learning with verifiable rewards (RLVR) on labeled data. We propose Continuous Test-Time Reinforcement Learning (C-TTRL), which uses a learned critic to reward policy updates on unlabeled test problems. At each iteration, verified responses generated by the updated policy on a separate, non-overlapping labeled reference set are used solely to recalibrate the critic. Across OLMo3 and Qwen3 models, C-TTRL continues to improve over the evaluated training horizon while representative self-rewarded baselines plateau. On AIME 2024, it raises Avg@16 from 33.0% to 51.1% for OLMo3-7B and from 42.3% to 65.8% for Qwen3-14B. Controlled comparisons support the respective roles of unlabeled test problems and critic recalibration. Code is available at https://anonymous.4open.science/r/C-TTRL-5B4F.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.