FedAttr: Towards Privacy-preserving Client-Level Attribution in Federated LLM Fine-tuning
Abstract
Data owners often need a way to verify whether a released corpus is used for unauthorized model training purposes by users. Training-data watermarking enables such verification via modifying the corpus so that models trained on it exhibit a detectable signal. In federated learning (FL), the global model aggregates updates from multiple clients, thus detecting the training-data watermark in the global model is feasible to check the corpus usage, but cannot identify which clients used it, which is the task we call client-level watermark attribution. Testing each client's locally updated model individually would reveal such attribution, but would also expose the client's update with privacy concern. Secure aggregation (SA) protects the individual updates via only exposing their aggregation but hides client-specific detection information which is needed for attribution, i.e. preventing the data owner from identifying which clients used the corpus. We propose FedAttr, a client-level watermark attribution protocol for FL with using the secure aggregates. For each target client, FedAttr contrasts aggregates from random subsets that include or exclude the client, yielding an unbiased update estimate with bounded variance. It then applies differential watermark scoring and combines evidence across communication rounds. We prove exponential attribution error bounds and bound the estimator's per-round mutual-information leakage by . In federated LoRA fine-tuning, FedAttr achieves 100% TPR at 0% FPR within five rounds across two watermark families and two aggregation strategies, outperforming all baselines, with 6.3% overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.