Rethinking Intermediate Supervision in Federated Learning with Adaptive Distillation
Abstract
Federated Learning (FL) enables multiple clients to collaboratively train a shared model while keeping their data local. Recent FL methods increasingly introduce supervision at intermediate layers, yet the effects of different forms of supervision in FL remain unclear. We examine Cross-Entropy (CE) and Knowledge Distillation (KD) for intermediate branch supervision in FL. While CE directly uses hard labels, KD provides richer guidance through the final classifier's soft predictions. Our analysis shows that this distinction becomes particularly important under federated local training, where CE is more sensitive to limited local data while KD provides more stable guidance. Motivated by these findings, we adopt KD for intermediate supervision. Its appropriate strength, however, can vary across federated settings and during training, so we propose adaptive KD weighting that adjusts the distillation strength to local training conditions. Experiments across multiple datasets and FL algorithms show that the proposed approach provides consistent performance across diverse federated settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.