acceptodds
Under review as a conference paper at ICLR 2027

WRONG VOTES, RIGHT GRADIENTS: WHY MAJORITY-VOTE REWARDS TRACK VERIFIABLE RL

Abstract

Reinforcement learning with self-generated labels, such as majority vote over a model's own samples, often recovers most of the gain of training with verifiable rewards. This is puzzling from the usual prompt-level view, in which majority vote only returns a model's modal answer and cannot add information about correctness. We argue that the relevant unit is the batch-level update. Because parameters are shared, a prompt's own gradient accounts for less than 1% of the change in that prompt's output at practical batch sizes, so a wrong label matters only to the extent that label errors share a direction across prompts. We decompose the majority-vote update into shrinkage and misdirection relative to the gold update, and show that misdirection vanishes with batch size whenever errors are incoherent across prompts, however often the labels are wrong. Measuring per-response gradients on five models, we find that majority-vote errors are as incoherent as a random-reward null, that the majority-vote update agrees with the gold update as closely as two independent gold updates agree with each other, and that its noise-corrected cosine with the gold update is 1 within error. Randomly flipping up to 30% of gold labels leaves the update direction intact, while the same fraction flipped by a shared rule destroys it, and rewards that favor a kind of response on every prompt, such as a learned reward model, fall measurably short. The equivalence holds at every checkpoint we examine but is local: gold-reward training still ends ahead on held-out majority accuracy. Error coherence, not error rate, determines when self-generated labels can stand in for verification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.