acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Intermediate Supervision in Federated Learning with Adaptive Distillation

Abstract

Federated Learning (FL) enables multiple clients to collaboratively train a shared model while keeping their data local. Recent FL methods increasingly introduce supervision at intermediate layers, yet the effects of different forms of supervision in FL remain unclear. We examine Cross-Entropy (CE) and Knowledge Distillation (KD) for intermediate branch supervision in FL. While CE directly uses hard labels, KD provides richer guidance through the final classifier's soft predictions. Our analysis shows that this distinction becomes particularly important under federated local training, where CE is more sensitive to limited local data while KD provides more stable guidance. Motivated by these findings, we adopt KD for intermediate supervision. Its appropriate strength, however, can vary across federated settings and during training, so we propose adaptive KD weighting that adjusts the distillation strength to local training conditions. Experiments across multiple datasets and FL algorithms show that the proposed approach provides consistent performance across diverse federated settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.