Two Momenta: A Training-Dynamics View of the EMA Teacher
Abstract
Self-distillation trains a student against a teacher that is an exponential moving average (EMA) of its own weights. Weight averaging is commonly interpreted as a filter: slow changes pass to the teacher largely intact, whereas fast changes are attenuated by an amount determined by the EMA momentum. In self-distillation, however, the filtered trajectory feeds back into the optimization process that generates it. We therefore analyze the EMA teacher explicitly as a filter and ask what it attenuates at each timescale. We answer that question on behavior rather than weights, from frequency-domain measurements of what both networks compute on a fixed set of images at every step. The student trajectory contains a dominant narrow oscillation, one to two orders of magnitude faster than anything the averaging can follow, so the teacher accumulates the student rather than averaging recent copies of it. That oscillation is a resonance of the student's other momentum, the optimizer's: a curvature fitted at one momentum predicts how its period shifts at three others, and the two momenta are coupled through the attenuation of this mode. Removing optimizer momentum at a matched effective step eliminates the identifiable resonance and sharply reduces the teacher's benefit. On CIFAR-100, removing the teacher costs 11.5 linear-probe points at the standard momentum and 0.7 without optimizer momentum. This reduction persists across three additional datasets, including ImageNet-1k, and at twice the training budget; removing the teacher also halves the model's forward FLOPs per training step.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.