Membership Inference Attacks on Language Models with Multiple Alignment Methods
Abstract
Large language models (LLMs) are fine-tuned by different methods on different data during their alignment for knowledge acquisition or preference optimization. The data used in this stage may leak sensitive information. In this work, we propose membership inference attacks for multiple and multi-stage LLM alignment methods. We introduce a extensible set of signals covering different aspects of the alignment pipeline to map the sample data to a multi-dimensional vector. Depending on the availability of shadow models, we propose shadow-free **Envelope (ENV)** and shadow-based **shadow-calibrated likelihood ratio (SLR)** to construct the final composite score for membership inference. We evaluate the framework on five model families (GPT-2, Llama-3.2-1B, Qwen2.5-1.5B-Instruct, Qwen2.5-3B-Instruct, and Llama-3.1-8B-Instruct) under two reward granularities(a continuous learned reward and a binary verifier) and multiple alignment pipelines. The extensive results indicate our proposed methods can achieve stronger average AUC than existing methods across all configurations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.