acceptodds
Under review as a conference paper at ICLR 2027

RAISE: Post-Training Membership Inference Mitigation via Model Editing

Abstract

Membership inference attacks (MIAs) pose privacy risks to fine-tuning data. After fine-tuning, a privacy audit may reveal that particular fine-tuning records carry strong membership signals. Yet applying training-time defenses at this point would require rerunning fine-tuning. We present a novel post-training approach that formulates MIA mitigation as a model-editing problem. Given selected member records, our approach, RAISE (eference-ligned nference-ignal diting), directly edits the fine-tuned checkpoint without rerunning fine-tuning. Our analysis shows that the membership signal is concentrated in a small number of token positions with unusually low losses relative to non-member references. RAISE localizes the corresponding spans and uses constrained checkpoint editing to align their losses with ranges derived from non-member references, while limiting changes to unrelated model behavior. Across six domains and ten MIAs, RAISE reduces mean attack AUC from 0.855 to 0.576 and mean TPR at 1% FPR from 25.1% to 2.8%. It also outperforms the evaluated post-training baselines, generalizes across five model configurations, and remains effective against adaptive attacks (AUC below 0.58), while retaining substantial fine-tuning utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.