acceptodds
Under review as a conference paper at ICLR 2027

Continuous Test-Time Training for Large Reasoning Models

Abstract

Scaling test-time computation improves the reasoning performance of large language models (LLMs) without improving the models themselves. Test-time training (TTT) instead updates model parameters on unlabeled test problems. Without ground-truth answers, existing methods often use agreement or confidence as proxy rewards, achieving early gains that can plateau or reverse with further training. We investigate whether TTT can sustain improvement beyond the plateau of reinforcement learning with verifiable rewards (RLVR) on labeled data. We propose Continuous Test-Time Reinforcement Learning (C-TTRL), which uses a learned critic to reward policy updates on unlabeled test problems. At each iteration, verified responses generated by the updated policy on a separate, non-overlapping labeled reference set are used solely to recalibrate the critic. Across OLMo3 and Qwen3 models, C-TTRL continues to improve over the evaluated training horizon while representative self-rewarded baselines plateau. On AIME 2024, it raises Avg@16 from 33.0% to 51.1% for OLMo3-7B and from 42.3% to 65.8% for Qwen3-14B. Controlled comparisons support the respective roles of unlabeled test problems and critic recalibration. Code is available at https://anonymous.4open.science/r/C-TTRL-5B4F.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.