acceptodds
Under review as a conference paper at ICLR 2027

Consensus Is Not Correctness: Reliability-Aware Self-Teaching for Unsupervised Self-Distillation

Abstract

Unsupervised on-policy self-distillation uses a model's own consensus to construct training signal, yet answer agreement does not certify the reasoning trace being imitated. We ask which quantities consensus actually certifies and introduce Reliability-Aware Self-Teaching (RAST), a control layer that separates answer confidence, group trace consistency, teacher-process reliability, and retained supervision. In process-auditable diagnosis logs, answer-correct traces can fail process audit in both Gemini 2.5 Pro and a dedicated Qwen3-4B process-gold audit subset, although process measurability is selective and the conditional rates are not population prevalence estimates. Answer confidence and teacher reliability specialize to different audit targets (AUROC 0.894 vs. 0.624 and 0.649 vs. 0.788), while joint thresholding traces a risk-coverage frontier below the audit-defined risk rate expected under random thinning of the same candidate pool. Correction-time selector rankings also disagree between answer and process objectives. A frozen five-seed LoRA study does not establish a multiplicity-corrected accuracy winner. A post-freeze three-seed MATH-500 downstream stress test likewise does not resolve a Majority@8 advantage for either self-distillation condition: the RAST-U-OPSD difference is -0.07 percentage points (95% paired two-way bootstrap CI [-1.40,+1.33]). The supported conclusion is narrower: self-distillation should monitor answer trust, process trust, and supervision coverage separately rather than collapse them into a single consensus score.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.