acceptodds
Under review as a conference paper at ICLR 2027

CoVER: Contrastive Self-Verification for Label-Free Reinforcement Learning

Abstract

A model that improves itself on unlabeled inputs would reduce its reliance on external supervision. Test-time reinforcement learning takes a step toward this goal by deriving rewards from the model's own predictions. However, agreement-based rewards only reflect the model’s answer distribution: they can sharpen an existing belief, but cannot tell whether that belief is correct. This creates a fundamental form of confirmation bias. A programmatic checker provides an external correctness signal without requiring labeled training examples, but it is often sparse, task-specific, or unavailable until a useful partial solution is produced. We propose CoVER (COntrastive VErification Reward), a hybrid label-free reward that combines programmatic feedback with relative self-verification. \mname first uses contrastive verification to distinguish promising candidate solutions and maintain a useful distribution of trajectories, then incorporates programmatic feedback to provide grounded credit for valid partial progress. This combination allows the model to exploit executable evidence when available while retaining a model-derived signal when programmatic feedback is insufficient. Across diverse reasoning tasks, CoVER consistently outperforms existing label-free rewards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.