Who Influenced This Behavior? Auditing Federated LLMs Without Revealing Client Updates
Abstract
Federated parameter-efficient fine-tuning (PEFT) can adapt language models to decentralized data while secure aggregation hides individual dense numerical learning updates. This confidentiality creates a post-hoc accountability problem: given an observed model behavior, how can an auditor rank the clients whose updates locally supported or opposed it? We introduce AudiTrace, a behavior-conditioned client-ranking method based on magnitude-free audit records. At selected checkpoints, each client locally records only the coordinate indices and signs of its Top- trainable-update entries; the dense update can remain on an independent secure-aggregation learning path. Given a differentiable audit objective, the auditor constructs an analogous query-gradient record and aggregates FedAvg-weighted signed overlap across checkpoints, yielding a directional ranking without reconstructing individual dense updates. We characterize exact mask-space geometry, derive conditional approximation bounds, and formalize a hidden-tail failure mode. Across diverse model, task, and partition settings, AudiTrace reaches dense-reference Spearman correlation –. Against exact leave-one-client-out retraining it attains Spearman and positive-effect AUROC from coordinate identities and signs alone, within Spearman of dense cosine. A signature retains Spearman with a 415.4 KiB audit record; repeated-sign membership inference peaks at AUROC versus chance. AudiTrace thus enables post-hoc auditing in federated PEFT while leaving individual dense updates on the secure-aggregation path.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.