acceptodds
Under review as a conference paper at ICLR 2027

Same Weights, Different Futures: State Identifiability in On-Policy Distillation

Abstract

On-policy distillation changes not only a student's predictions but also the states on which it receives future supervision. Thus, two students that currently appear identical may still have different probabilities of recovering a missing behavior within a finite query budget. Specifically, we study which representation of training state determines this recoverability. To formalize this, we define budgeted recoverability on a transition-sufficient state and show that an observable summary is sufficient exactly when all full states sharing that summary have the same recovery probability. The resulting observation-fiber diameter gives a sharp minimax lower bound on prediction error from the summary alone. We then construct two positive-probability histories of one finite OPD process that have identical observations but different recovery laws, and show that Adam can preserve such differences in optimizer state even when weights are identical. We further characterize scalar recovery regimes and decompose the value of finite-query intervention into state access and the signed reliability of teacher feedback. Empirically, common-future Transformer experiments realize the exact-weight separation, while optimizer cross-swaps test the same mechanism under ordinary training in a procedural task and Qwen3.5 reasoning on GSM8K. Prospective validation indicates that recovery-gap magnitude transfers more reliably than state ordering. Together, these results show that recoverability is determined by the state consumed by future transitions, not by current weights or predictions alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.