acceptodds
Under review as a conference paper at ICLR 2027

FedDataLeak: Detecting Data Leakage from Federated LLM Updates

Abstract

Federated fine-tuning allows large language models (LLMs) to learn from decentralized, sensitive data without sharing raw text, yet the exchanged model updates can still reveal private information. We present FedDataLeak, a model-agnostic framework for auditing leakage from federated LLM updates using model weights before and after each communication round. Its primary signal, Feature Reconstruction Fidelity (FRF), measures how well a plausible input can reproduce an observed update. To avoid costly reconstruction at every round, FedDataLeak combines FRF with two lightweight representation-shift signals, neuron polysemanticity and activation sparsity, and triggers deeper analysis only for suspicious updates. We evaluate the framework across heterogeneous federated settings, including controlled client partitions, synthetic canaries, and aggregate-only visibility, and test whether its scores identify leakage-prone updates and align with downstream reconstruction and membership evidence. Across diverse models and datasets, FedDataLeak achieves strong leakage detection with audit overhead lower than competitive baselines, demonstrating that federated updates can serve as actionable privacy signals for safe LLM training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.