acceptodds
Under review as a conference paper at ICLR 2027

Training-Sufficient Supervision: What Distillation Caches Must Preserve

Abstract

A distillation cache can be built to reconstruct the teacher's prediction, or to reproduce the training a student would have performed. These are different requirements. For cross-entropy under a frozen affine head the teacher enters the update only through , so targets differing elsewhere on the simplex induce identical gradients. We call a representation training-sufficient when it preserves the gradient field over the whole declared parameter domain rather than at one iterate, and characterize it: a fixed linear cache serves every target exactly when the reachable space of centered logit changes lies in its column space, and the least number of coordinates it can use is the dimension of that space. Two constructions follow, one for a frozen head and one for output directions opened on a schedule, the second identifying what must be appended when training reaches directions the cache does not cover. Controlled interventions confirm the predicted separation and its repair. Trained at a matched byte budget against renormalized Top-k and a low-rank reconstruction of the teacher distribution, the two exact representations hold gradient error four to five orders of magnitude lower along the whole trajectory, and store the supervision in under 5% of a full target's bytes, which on one pair yields a measured per-step rate. Sufficiency remains a property of the declared loss and training family, not of the teacher alone: reversing the divergence, with the student untouched, raises the requirement to coordinates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.