acceptodds
Under review as a conference paper at ICLR 2027

Towards Probabilistic Unlearning Attacks on Large Language Models

Abstract

Evaluating large language model (LLM) unlearning requires determining whether target information remains accessible after its removal. However, unlearning is commonly evaluated and attacked through single responses, even though target information may remain exposed under repeated sampling. This probabilistic exposure creates an additional attack surface beyond single-response relearning attacks. We introduce the first probabilistic unlearning attack that directly optimizes target-fact disclosure over sampled responses. We derive finite-sample confidence bounds, characterize its idealized dynamics, and establish conditions for greater disclosure than single-reference supervised fine-tuning (SFT). Across Harry Potter QA, TOFU, and WMDP, we show that single-response metrics can miss target exposure, GRPO recovers unlearned information more effectively than state-of-the-art SFT, and it substantially compromises RULE-NPO, a recent method designed to resist probabilistic leakage. Our results reveal probabilistic outputs as an important attack surface for LLM unlearning. Anonymous source code: https://anonymous.4open.science/r/relearning_attack_grpo-321B

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.