acceptodds
Under review as a conference paper at ICLR 2027

Beyond Self-Consensus: Mitigating Self-Reinforcing Exploration Bias in Test-Time Reinforcement Learning

Abstract

Test-time reinforcement learning (TTRL) enables language models to improve reasoning under unlabeled test-time settings through self-generated supervision. Yet this supervision is intrinsically coupled to the model’s own exploration distribution, creating an endogenous feedback process distinct from conventional reinforcement learning with externally grounded rewards. We identify this feedback as *self-reinforcing exploration bias*, characterized by the amplification of existing policy preferences through endogenous pseudo-supervision and the progressive concentration of subsequent exploration around a restricted set of solution modes. To mitigate this bias, we introduce a two-dimensional structured exploration framework that expands reasoning-space coverage through multiple model-driven reasoning views and adapts exploration depth to instance-specific search demand. The resulting framework increases access to alternative solution trajectories while allocating test-time computation according to heterogeneous exploration requirements. Across multiple reasoning benchmarks and model scales, our approach consistently improves TTRL performance while reducing sampling cost by more than 30%. Our analysis shows that the resulting limitation is not solely a consequence of noisy pseudo-rewards or unstable optimization, but also reflects insufficient control over the solution space exposed during adaptation. These findings establish exploration as a quantitatively actionable dimension of TTRL, providing a principled route toward mitigating self-reinforcing adaptation under self-generated supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.