acceptodds
Under review as a conference paper at ICLR 2027

Diana: Calibrating the Per-Token Null for Membership Inference on Fine-Tuned LLMs

Abstract

Membership inference attacks on language models detect whether a specific text was included in the training data. Fine-tuning on a domain complicates this task: an increase in a token’s log-probability over the pre-trained model, which we call its gain, may reflect domain-level learning rather than memorization of the target sequence itself, inflating false positives in reference-based attacks. We introduce *Diana* (Distribution-Informed Attacks with Null Adjustment), a plug-and-play calibration that isolates memorization from domain adaptation. Our two key observations are: (i) for non-members, the target model’s own next-token distribution approximates the domain distribution. This motivates centering: subtracting the target-weighted mean gain over possible next tokens from the observed gain. (ii) Under our modeling assumptions, the centered gain has zero conditional mean for non-members and positive conditional mean for members. To account for differences in background variability, *Diana* divides the centered gain by the regularized gain variance to obtain a calibrated gain. Hence, *Diana* requires only the log-probabilities of the target and pre-trained models—no auxiliary data or additional training—and integrates with any reference-based attacks, which then aggregate the calibrated gains under their original rules. Across six corpora, *Ratio**Diana* achieves the mean true-positive rate of *Ratio* at a false-positive rate and outperforms all evaluated prior attacks on every corpus.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.