AliasGuard: Surrogate-Based Defense Against Passive Inference in Split Language Model Fine-Tuning
Abstract
Split learning (SL) is often treated as a privacy-preserving way to fine-tune language models (LMs) without exposing clients' private training tokens. In split-based LM fine-tuning (SL-FT), however, an honest-but-curious (passive) server can leverage smashed data in the forward pass and gradients in the backward pass to infer these private tokens in training sequences. Under the passive inference, current defenses face a trade-off: they cannot protect these private tokens while retaining training utility. To mitigate this trade-off, we present AliasGuard, a surrogate-based defense. Specifically, AliasGuard generates surrogates from the unprotected context of a sequence and uses these surrogates to construct the training sequences and one-hot vectors for computing smashed data and gradients. We prove that AliasGuard's smashed data and gradients provide no additional information about private tokens beyond the unprotected context. We also bound changes in smashed data at unprotected-token positions. We evaluate AliasGuard under three corpora and three models. Under adaptive inference attacks, AliasGuard reduces private-token recovery by 67.4-99.6% relative to the lowest baseline rate, while retaining much of No Defense baseline's task utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.