acceptodds
Under review as a conference paper at ICLR 2027

ATTENMIA: READING MEMBERSHIP OFF THE ATTENTION MAPS OF LANGUAGE MODELS

Abstract

Large language models (LLMs) can reveal through attacks such as Membership Inference Attacks (MIAs) whether a sequence appeared in their training data, creating privacy risks. Many existing MIAs for LLMs derive their signal from model outputs. Motivated by recent research showing the important role attention plays in information flow within a model, in this paper, we instead ask whether membership signatures are visible in the model’s attention patterns. We introduce AttenMIA, a white-box MIA with two variants: Transitional features capture cross-layer attention changes, while Perturbation features measure attention shifts under input edits. Each variant uses a lightweight supervised classifier to convert its features into membership scores. We evaluate on MIMIR, a benchmark constructed to reduce known distributional pitfalls in membership evaluation. Across seven domains and three Pythia model sizes, AttenMIA (Perturbation) achieves the highest ROC AUC in 20 of 21 settings, with domain-averaged AUCs of 0.83–0.85, compared with 0.71 for the strongest evaluated baseline. When each domain is held out from attack training and selection, AttenMIA (Transitional) performs best among the evaluated methods across all three model sizes, averaging 0.73 ROC AUC and 26.7% TPR@1%FPR across domains and models. Under the same protocol, a supervised hidden-state control achieves an average AUC of 0.54, showing that the advantage arises from attention features. GPT-2 fine-tuning experiment shows that cross-layer transitional attention features, and perturbation-response features have different attention gradients for members and non-members. These results substantially improve on SoTA membership inference attacks, and demonstrate that self-attention provides a useful signal for training-membership inference that transfers across the evaluated domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.