Temporal Loss Signatures for Data Pruning in Federated Learning with Noisy Labels
Abstract
Federated learning enables collaborative model training without sharing raw data, but heterogeneity and label noise can degrade performance. Data pruning is difficult because high loss conflates mislabeled examples with hard but informative ones, while static selections become stale as the global model evolves. We introduce the Loss Signature Score (LSS), a dynamic, client-local criterion that combines sliding-window mean loss with loss gain. Mean loss captures persistent difficulty, whereas gain favors hard examples that show learning progress. Periodic re-scoring, class-aware retention quotas, and FedProx enable adaptive pruning without estimating the noise rate. The framework adds one full-local-data inference pass per round but requires no scoring-related backward pass or sample-level communication. Experiments on CIFAR-10/100 and CIFAR-10N/100N cover non-IID partitions, synthetic and human annotation noise, and nominal pruning rates from 10% to 70%, with the largest benefits under stronger noise and aggressive pruning. On CIFAR-100 with 20% symmetric noise and a 50% pruning target, LSS achieves 65.38% test accuracy, exceeding the next-best baseline by 4.29 percentage points. On CIFAR-100N with 40.2% human-annotated label noise, it outperforms all evaluated data-pruning baselines at every tested pruning rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.