Reflection-Conditioned Teacher–Student Test-Time Reinforcement Learning
Abstract
Test-time reinforcement learning aims to adapt large language models on unlabeled queries using only self-generated reasoning trajectories. This setting is attractive when ground-truth labels, executable verifiers, or trusted reward models are unavailable, but it also makes the learning signal fragile. Existing self-consistency-based objectives often use the same query-only rollouts both to define pseudo-rewards and to receive updates, which can turn unreliable agreement into overconfident supervision. Moreover, standard policy-gradient objectives apply a trajectory-level advantage to all tokens, despite the fact that only a subset of tokens is likely to determine the final answer. We propose RCTS, a reflection-conditioned teacher–student framework for label-free test-time reasoning adaptation. RCTS uses a single policy snapshot in two contextual roles: query-only student rollouts reveal the model's initial behavior, while reflection-conditioned teacher rollouts provide a structured reference. From the teacher rollouts, RCTS constructs a teacher-weighted answer distribution that assigns soft consistency scores to student trajectories. It further introduces Cross-Group Token Divergence to select student tokens that most differ from the reflection-conditioned teacher distribution. The final update applies mask-selective GRPO only to selected query-only student tokens, regularized by a KL anchor to the base policy. Experiments across mathematical reasoning and knowledge-intensive QA show that RCTS consistently improves the frozen models and is competitive with representative label-free test-time adaptation baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.