acceptodds
Under review as a conference paper at ICLR 2027

Privacy leakage or distribution shift? Post-hoc Privacy Auditing of foundation models

Abstract

Membership inference attacks (MIA) provide a theoretically grounded basis for auditing privacy leakage from machine learning models, but applying it in a post-hoc manner to pre-trained large language models (LLMs) is challenging. Indeed, MIAs quantify leakage by distinguishing data used in training (members) from non-members from the same distribution. However, in practice, auditors have to collect non-members from different sources, resulting in distribution shift that biases privacy leakage measurements. To solve this issue, we propose two general extensions to recently proposed debiased post-hoc privacy auditing techniques to make them practical for LLMs. First, we reduce the non-member distribution shift by rewriting non-members in the style of member data and filter members and non-members to focus on regions of greater empirical overlap. Second, we strengthen MIAs on LLMs by adding paraphrase-related features that compare model behaviour on suspected members with model behaviour on semantically similar paraphrases. A thorough experimental evaluation on text and vision-language benchmarks with temporal, formating and collection shifts, demonstrates that our approaches decrease distribution shift while retaining detectable membership signals, thus yielding meaningful privacy leakage estimates. Our results also show that distribution shifts resulting from naturally collected non-members does not preclude practical post-hoc privacy auditing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.