acceptodds
Under review as a conference paper at ICLR 2027

AliasGuard: Surrogate-Based Defense Against Passive Inference in Split Language Model Fine-Tuning

Abstract

Split learning (SL) is often treated as a privacy-preserving way to fine-tune language models (LMs) without exposing clients' private training tokens. In split-based LM fine-tuning (SL-FT), however, an honest-but-curious (passive) server can leverage smashed data in the forward pass and gradients in the backward pass to infer these private tokens in training sequences. Under the passive inference, current defenses face a trade-off: they cannot protect these private tokens while retaining training utility. To mitigate this trade-off, we present AliasGuard, a surrogate-based defense. Specifically, AliasGuard generates surrogates from the unprotected context of a sequence and uses these surrogates to construct the training sequences and one-hot vectors for computing smashed data and gradients. We prove that AliasGuard's smashed data and gradients provide no additional information about private tokens beyond the unprotected context. We also bound changes in smashed data at unprotected-token positions. We evaluate AliasGuard under three corpora and three models. Under adaptive inference attacks, AliasGuard reduces private-token recovery by 67.4-99.6% relative to the lowest baseline rate, while retaining much of No Defense baseline's task utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.