acceptodds
Under review as a conference paper at ICLR 2027

Fair Credit Across Correct Modes: Mitigating Correct-Path Coverage Collapse in On-Policy Self-Distillation for Language Models

Abstract

On-policy self-distillation (OPSD) is a promising approach for improving reasoning in large language models through dense supervision from a context-augmented version of the model. However, OPSD with sampled demonstrations (OPSD-SD) can achieve strong pass@1 accuracy while showing limited gains from additional rollouts, a phenomenon we call correct-path coverage collapse. Beyond general mode-seeking tendencies, demonstration conditioning in OPSD-SD can exacerbate this collapse through pointwise conditional mutual information (PCMI)-induced sharpening, further favoring already-probable correct solutions. To this end, we propose \ours, an OPSD framework that combines PCMI Debiasing (PD) and Diversity-Aware Exploration (DAE). PD corrects demonstration-dependent preferences among equally correct responses. Under idealized conditions, we show that this correction exactly cancels the additional PCMI-induced sharpening between correct responses. DAE then assigns credit according to each correct path’s marginal contribution to diversity, encouraging underrepresented paths. Experiments across multiple datasets demonstrate that \ours mitigates correct-path coverage collapse while outperforming strong baselines on both Pass@1 and Pass@.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.