From Clients to Samples: Tracing Data Influence in Federated Learning
Abstract
Understanding how individual training data shape model behavior is essential for error diagnosis and data curation, particularly in medical applications. In Federated Learning (FL), tracing training-data influence is particularly challenging because data remain decentralized, and requires separating the effects of decentralized client updates from those of individual local samples. Existing methods primarily focus on quantifying client-level contributions, offering limited insight into the individual samples that shape their local models. Such aggregate assessments can mask samples with heterogeneous or opposing effects on model behavior, making targeted data inspection difficult. To address these challenges, we propose Federated Hierarchical Influence Tracing (FedHIT), an attribution framework for medical FL that traces influence at two complementary levels. Level 1 attributes query-specific global model behavior to individual clients' local updates, identifying the clients most responsible for performance on a given query. Level 2, applied within a client identified at Level 1, attributes that client's local model behavior to its own training samples, identifying which examples drive its local model. This decomposition connects global behavior to responsible clients and, separately, a client's local behavior to responsible samples, without requiring influence to be traced end-to-end from individual samples to the global model. We validate FedHIT on three benchmark datasets, demonstrating that it reliably identifies influential clients and samples at each level, enabling targeted inspection of the data underlying client- and global-model behavior in federated medical settings. By restricting sample-level analysis to selected clients, FedHIT avoids exhaustive attribution over all decentralized samples, providing a computationally efficient approach to targeted data inspection and curation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.