CoVER: Contrastive Self-Verification for Label-Free Reinforcement Learning
Abstract
A model that improves itself on unlabeled inputs would reduce its reliance on external supervision. Test-time reinforcement learning takes a step toward this goal by deriving rewards from the model's own predictions. However, agreement-based rewards only reflect the model’s answer distribution: they can sharpen an existing belief, but cannot tell whether that belief is correct. This creates a fundamental form of confirmation bias. A programmatic checker provides an external correctness signal without requiring labeled training examples, but it is often sparse, task-specific, or unavailable until a useful partial solution is produced. We propose CoVER (COntrastive VErification Reward), a hybrid label-free reward that combines programmatic feedback with relative self-verification. \mname first uses contrastive verification to distinguish promising candidate solutions and maintain a useful distribution of trajectories, then incorporates programmatic feedback to provide grounded credit for valid partial progress. This combination allows the model to exploit executable evidence when available while retaining a model-derived signal when programmatic feedback is insufficient. Across diverse reasoning tasks, CoVER consistently outperforms existing label-free rewards.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.