acceptodds
Under review as a conference paper at ICLR 2027

Collective Wisdom: On-Policy Self-Distillation with Synergistic Multi-Perspective Teacher Policy Fusion

Abstract

In On-Policy Self-Distillation (OPSD), the student receives dense token-level supervisory signals from a teacher branch conditioned on privileged information, thereby improving its reasoning ability without relying on a stronger external teacher. However, a single teacher perspective can introduce bias and incomplete coverage in the supervisory signals, causing student updates to rely excessively on local supervisory signals and leading to performance instability or even collapse during the distillation process. Although incorporating multi-perspective teachers can broaden the sources of supervisory signals, conflicts among the supervisory signals provided by different teachers can make inadequate policy fusion susceptible to performance fluctuations during the distillation process. To address these challenges, we propose Collective Wisdom (CW), a synergistic multi-perspective teacher policy fusion method for OPSD. Specifically, we first take the union of the top- vocabulary from each teacher to construct a joint policy space and normalize the individual teacher policies within this space to fuse them. We then realign the shared top- vocabulary based on the fused policy to distill a consensus policy, which is employed to optimize the student policy. This design enables the synergistic fusion and stable distillation of multi-perspective teacher policies. Furthermore, experiments on models of different sizes and architectures across four mathematical reasoning benchmarks show that our method outperforms representative baselines in stabilizing distillation and mitigating performance collapse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.