PLACEBU: Placebo-Controlled Retain-Free LLM Unlearning
Abstract
LLM unlearning aims to remove the influence of designated training data while preserving the model's remaining capabilities, with applications to privacy, copyright compliance, the removal of sensitive knowledge, and many more. To preserve model utility after unlearning, many existing methods rely on accessible retain data, typically using a small retain set as a proxy of the much larger original training corpus. However, such data can be unavailable in practice, and a randomly sampled retain set may not provide the most informative signal for determining what should be preserved. In this paper, we show that sampled retain data are not the only source of preservation supervision: the forget set itself contains targeted information about what should remain unchanged. Each forget example pinpoints the knowledge whose training influence is to be removed, while structured perturbations around that example probe the local predictive structure, separating forget-specific effects from shared effects that should be preserved. Based on this insight, we propose PLACEBU, which uses matched target-placebo controls to estimate this boundary, to achieve unlearning without access to retain examples. Experiments on TOFU and MUSE demonstrate competitive forgetting–utility trade-offs against both retain-based and retain-free baselines. On MUSE-Books, PLACEBU reduces knowledge memorization by 84.7% relative to a strong baseline while maintaining competitive utility preservation. Our code is available at https://anonymous.4open.science/r/PLACE-96A3/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.