FedTeas: Federated Learning via Teacher Distillation and Sharpness-Aware Minimization
Abstract
Data heterogeneity remains a central challenge in Federated Learning (FL), inducing client drift and often steering optimization toward sharp solutions with poor generalization. While Sharpness-Aware Minimization (SAM) can improve flatness and robustness, locally computed SAM perturbations may become misaligned with the global objective under heterogeneous data. To address this, we introduce FedTeas, a family of three methods: vanilla FedTeas, FedTeasC, and FedTeasM that use global knowledge distillation (KD) as directional guidance for local sharpness-aware optimization. By aligning local perturbations with global predictive knowledge, FedTeas mitigates client drift and yields more consistent, robust global models across communication rounds. Rather than applying KD as a conventional parameter-wise regularizer, FedTeas injects distillation gradients directly into the SAM perturbation step, aligning local perturbations with global predictive knowledge. We prove that vanilla FedTeas retains the convergence rate of FedSAM (seminal approach adapting SAM to FL), while the remaining family members preserve the convergence guarantees of their corresponding optimization frameworks. We establish a -conditioned generalization bound that characterizes the impact of KD-guided perturbation directions on federated generalization. Extensive experiments across heterogeneous federated benchmarks show that FedTeas consistently improves accuracy, robustness, and generalization over state-of-the-art federated optimization and sharpness-aware baselines. Code is available at https://anonymous.4open.science/r/FedTeas-17D1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.