The Wisdom of Wrong Crowds: Learning from Disagreement Among Failed Rollouts
Abstract
Reinforcement learning with verifiable rewards (RLVR) trains language models from verifiable outcomes, but with binary rewards, the outcome-based advantage in group-relative methods vanishes when every rollout for a problem is incorrect. Many recent self-distillation methods provide token-level supervision by modulating verifier-derived advantages, introducing external correct information, or constructing reference-free signals for failed rollouts. However, relations among failed rollouts within an all-wrong group remain underexplored as a source of token-level supervision. We observe that all-wrong groups need not be information-free: failed rollouts can make different mistakes, and disagreement in their answers provides a coarse cue to differences in their reasoning. We introduce **SibCredit**, which converts these answer relations into selective token-level negative credit by re-scoring a target's sampled tokens under contexts constructed from other failed rollouts in the same group. Different-answer rollouts provide candidate negative evidence, but their opposition can also reflect shared conditioning effects, such as shifts in reasoning style. Same-answer rollouts serve as a control, retaining penalties only when different-answer opposition sufficiently exceeds same-answer opposition. Because every conditioning rollout is incorrect, SibCredit uses only induced probability decreases as candidate negative evidence and discards probability increases. The resulting token-level signal is constructed entirely from the current all-wrong group, while other groups use standard GRPO. Experiments on Qwen3-1.7B and Qwen3-4B across mathematical reasoning and code generation show that SibCredit consistently improves over GRPO and recent self-distillation baselines. Token-level analysis further shows that same-answer calibration retains a larger fraction of opposition strength on annotated error tokens than on epistemic expressions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.