acceptodds
Under review as a conference paper at ICLR 2027

Privacy-Preserving Federated Fine-Tuning of Language Models under Non-IID Data

Abstract

Differentially private federated fine-tuning enables collaborative adaptation of language models while protecting sensitive training records, but its utility is challenged by client heterogeneity and privacy-induced noise. Control variates mitigate client drift, yet noisy updates also perturb local trajectories and subsequent control states. We propose LSSCA, a federated fine-tuning framework that combines stochastic controlled averaging with fixed random subspace filtering. Constructed independently of private data without requiring public reference data, the filter preserves reference-subspace components and attenuates orthogonal components of the complete control-corrected noisy update. An exact signal–noise decomposition characterizes the trade-off between noise reduction and signal distortion, without assuming that the random subspace aligns with useful gradients. Under the stated assumptions, including inactive clipping, we establish an non-convex stationarity rate at a fixed noise scale without an additional gradient-dissimilarity assumption. The analysis explicitly accounts for noise propagation through local trajectories and control states, retaining the costs of filtering in the convergence constants. Experiments across NLP benchmarks, privacy budgets, and client heterogeneity levels demonstrate improved privacy–utility trade-offs over the evaluated DP-FL baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.