acceptodds
Under review as a conference paper at ICLR 2027

Where to Verify, Where to Explore: Test-Time Reinforcement Learning with Confidence-Conditioned Verification

Abstract

Test-time reinforcement learning (TTRL) has emerged as a promising paradigm for enhancing the complex reasoning abilities of large language models (LLMs) in a completely label-free manner. Despite existing studies focusing on Pass@1 performance, optimizing Pass@k remains under-explored yet critical in label-free settings, which measures generation coverage for sustained exploration. Optimizing Pass@k in label-free setting is highly non-trivial, as directly applying the Pass@k advantage designs effective for RLVR yields unsatisfactory performance. We show that this transfer utilizes samples uniformly across confidence regions, although samples in different regions differ in both exploration value and label reliability. Low-confidence samples carry the strongest exploration signal, but their mostly incorrect majority labels push exploration in the wrong direction; high-confidence samples have reliable labels, but their exploration value is nearly exhausted and their trajectory diversity collapses. To overcome these hurdles, we propose TTRL-CoCoV (Test-Time Reinforcement Learning with Confidence-Conditioned Verification), a novel confidence-adaptive framework that expands Pass@k coverage and improves Pass@1 by conditioning the operation applied to each sample on its confidence. For high-confidence samples, it uses their reliable pseudo-labels to supervise the verifier and applies a length-diversity reward to prevent diversity collapse; for low-confidence samples, it delegates pseudo-label selection to the verifier to filter out incorrect pseudo-labels; and for medium-confidence samples, it bypasses verification entirely. In this way, reliable high-confidence supervision improves the verifier, which in turn corrects low-confidence pseudo-labels, allowing generation and verification to co-evolve within a single model. Extensive experiments show that TTRL-CoCoV outperforms competing label-free methods across six benchmarks, with average absolute gains of 10.42 points in Pass@1 and 20.25 points in Pass@16 over TTRL. Anonymous code repository for TTRL-CoCoV: https://anonymous.4open.science/r/TTRL-CoCoV-C5BC.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.