acceptodds
Under review as a conference paper at ICLR 2027

Function-Space Temporal Self-Distillation

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) improves reasoning in language models, but scalar outcomes provide only coarse supervision over response trajectories. Self-distillation offers denser token-level guidance by comparing an unassisted policy with a feedback-informed self-teacher; however, because the teacher evolves alongside the student, its targets can become unstable, motivating additional regularization. We introduce Function-Space Temporal Self-Distillation (FTSD), a single-policy framework that combines current privileged adaptation with function-space temporal consolidation. FTSD stores sparse historical next-token distributions as fixed temporal targets, providing temporal separation without maintaining a persistent teacher model. It further incorporates privileged disagreement correction to suppress locally overvalued actions, reliability-adaptive grounding to strengthen reference supervision when self-generated evidence repeatedly fails, and quality- and coverage-aware temporal memory to retain verified behaviors and improve privileged demonstrations. Across reasoning and tool-use benchmarks, FTSD achieves the highest average accuracy under a 5-hour training budget, outperforming the strongest baseline by approximately 4 percentage points on average while reducing peak memory usage from about 70 GB to 56 GB on Olmo3-7B-Instruct.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.