NIFT: NOISE-INVARIANT FINE-TUNING FOR PRIVACY- PRESERVING LLMS
Abstract
Privacy-preserving fine-tuning of large language models is appealing for many important applications, but remains challenging in practice. Adding differentially private (DP) noise to input representations can protect private data, but this noise is amplified by deep Transformer layers and sharply degrades performance. Prior work has primarily addressed this problem through noise-correction mechanisms that attempt to recover useful representations after amplification has already oc- curred, but these approaches still leave a substantial performance gap compared with the corresponding no-noise setting. In this paper, we propose a two-stage training framework that instead prepares the backbone to be less sensitive to the agreed DP noise distribution before downstream task fine-tuning. First, we train a LoRA adapter with a label-free pairwise consistency objective over multiple noisy embedding copies, encouraging the backbone’s hidden states to remain stable under DP perturbations. Second, we fine-tune a lightweight client-side ladder side network that fuses local features with selected hidden states returned by the robust server-side backbone. This separation of noise-invariance preparation from task adaptation reduces the burden on the client model, improves robustness un- der privacy-preserving noise. Our experiments show that NIFT outperforms the state-of-the-art HiddenEcho across multiple datasets and privacy budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.