FedDataLeak: Detecting Data Leakage from Federated LLM Updates
Abstract
Federated fine-tuning allows large language models (LLMs) to learn from decentralized, sensitive data without sharing raw text, yet the exchanged model updates can still reveal private information. We present FedDataLeak, a model-agnostic framework for auditing leakage from federated LLM updates using model weights before and after each communication round. Its primary signal, Feature Reconstruction Fidelity (FRF), measures how well a plausible input can reproduce an observed update. To avoid costly reconstruction at every round, FedDataLeak combines FRF with two lightweight representation-shift signals, neuron polysemanticity and activation sparsity, and triggers deeper analysis only for suspicious updates. We evaluate the framework across heterogeneous federated settings, including controlled client partitions, synthetic canaries, and aggregate-only visibility, and test whether its scores identify leakage-prone updates and align with downstream reconstruction and membership evidence. Across diverse models and datasets, FedDataLeak achieves strong leakage detection with audit overhead lower than competitive baselines, demonstrating that federated updates can serve as actionable privacy signals for safe LLM training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.