acceptodds
Under review as a conference paper at ICLR 2027

Beyond Output-Level Guidance: Leveraging Internal Exploration for LLM Self-Distillation

Abstract

On-policy self-distillation (OPSD) provides dense token-level supervision for language model reasoning, but conditioning the self-teacher on privileged information may suppress the exploration needed by the student at inference time. By tracing middle-layer distributions throughout post-training from a interpretability perspective, we characterize how different training paradigms reshape the model's internal exploration–exploitation dynamics. In contrast to GRPO, OPSD progressively contracts latent exploration, and this contraction coincides with late-stage reasoning degradation. Motivated by this observation, we introduce retained exploration, a token-level signal that captures whether semantically diverse exploration persists across layers and identifies structurally important positions that shape subsequent reasoning trajectories. Based on this signal, we propose Retained-Exploration-Guided Self-Distillation (RESD), which preserves exploration-supported teacher guidance while counteracting potentially over-concentrated preferences at low-retained-exploration positions. Across three models and five mathematical reasoning benchmarks, RESD consistently outperforms strong RLVR and self-distillation baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.