Consensus as Privileged Context for Label-Free Self-Distillation
Abstract
Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large language models without labels, and a growing family of methods converts this consensus signal into training supervision. However, existing approaches use consensus only in restricted forms: as a filter that selects solutions for fine-tuning, as a preference between answers, or as a scalar reward for reinforcement learning, discarding most of the information that the agreeing solutions contain. We present CANON (Consensus-ANchored self-distillatiON), a label-free training method that turns consensus into dense, token-level supervision. For each unlabeled prompt, CANON samples multiple solutions, extracts the majority answer, and conditions a frozen snapshot of the model on a solution that reaches it; this consensus-anchored teacher then supervises the model on its own rollouts at every token. Experiments on mathematical and scientific reasoning benchmarks show that CANON improves pass@1 by up to 15 points, outperforming label-free reinforcement learning by 6 points at 15% of its compute and approaching a teacher conditioned on gold solutions. Trained on unlabeled data, it transfers to held-out benchmarks, matching training methods that use gold labels. Analysis attributes most of the improvement to the consensus being distilled into the weights and finds no significant narrowing of coverage. On the mathematics benchmarks, CANON's pass@ advantage over the base model persists at every sampling budget measured, reaching 512 samples per prompt, and its majority vote itself becomes more accurate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.