acceptodds
Under review as a conference paper at ICLR 2027

Learning to Detect Training Data via Active Reconstruction

Abstract

Detecting LLM training data is generally framed as a membership inference attack (MIA) problem. However, most MIAs operate passively on fixed model weights, using log-likelihoods or text generations. In this work, we introduce **Active Data Reconstruction Attack** (ADRA), a family of MIAs that actively induces a model to reconstruct a given text through training. We hypothesize that training data are *more reconstructible* than non-members, and the difference in their reconstructibility can be exploited for membership inference. Motivated by findings that reinforcement learning (RL) sharpens behaviors already encoded in weights, we leverage on-policy RL to actively elicit data reconstruction by finetuning a policy initialized from the target model. To effectively use RL for MIA, we design reconstruction metrics and contrastive rewards. The resulting algorithms, ADRA and its adaptive variant ADRA+, improve both reconstruction and detection given a pool of candidate data. Across pre-training, post-training, and distillation, our best method improves over the strongest available loss-based baseline by 11.2 AUROC points on average across 14 evaluation settings, and the strongest of all evaluated baselines by 6.8 points. As a highlight, ADRA+ improves over Min-K%++ by 18.8 points on BookMIA and 7.6 points on AIME.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.