acceptodds
Under review as a conference paper at ICLR 2027

FedAttr: Towards Privacy-preserving Client-Level Attribution in Federated LLM Fine-tuning

Abstract

Data owners often need a way to verify whether a released corpus is used for unauthorized model training purposes by users. Training-data watermarking enables such verification via modifying the corpus so that models trained on it exhibit a detectable signal. In federated learning (FL), the global model aggregates updates from multiple clients, thus detecting the training-data watermark in the global model is feasible to check the corpus usage, but cannot identify which clients used it, which is the task we call client-level watermark attribution. Testing each client's locally updated model individually would reveal such attribution, but would also expose the client's update with privacy concern. Secure aggregation (SA) protects the individual updates via only exposing their aggregation but hides client-specific detection information which is needed for attribution, i.e. preventing the data owner from identifying which clients used the corpus. We propose FedAttr, a client-level watermark attribution protocol for FL with using the secure aggregates. For each target client, FedAttr contrasts aggregates from random subsets that include or exclude the client, yielding an unbiased update estimate with bounded variance. It then applies differential watermark scoring and combines evidence across communication rounds. We prove exponential attribution error bounds and bound the estimator's per-round mutual-information leakage by . In federated LoRA fine-tuning, FedAttr achieves 100% TPR at 0% FPR within five rounds across two watermark families and two aggregation strategies, outperforming all baselines, with 6.3% overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.