Improved Convergence and Generalization Analysis for Online Knowledge Distillation
Abstract
Online knowledge distillation enables distributed models to learn collaboratively without requiring a pretrained teacher. However, the evolving and mutually dependent distillation targets complicate both convergence and generalization analysis. We study online distillation over decentralized networks with heterogeneous model architectures and personalized local distributions. We establish a finite-round convergence bound by controlling objective drift through the neighbors' local gradients and stochastic noise. This analysis does not rely on uniform bounds on gradients or objective drift used in existing analyses and yields tighter guarantees in wide-range parameter regimes. By combining our convergence and generalization analyses, we bound the personalized test risks of the local models produced after a finite number of training rounds. The resulting bound explicitly characterizes the roles of distribution discrepancies, dataset sizes, the trade-off between original supervised and distillation losses, and ensemble weights to construct teachers, including how these effects propagate through the network. This explicit parameter dependence enables tunable-parameter optimization based on a surrogate of the generalization bound. Using ensemble weigh optimization as example, our experiments demonstrate that this approach can substantially improve average personalized test accuracy across a range of heterogeneous settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.