Mode-Avoiding Unlearning via Negative-Order Rényi Divergence
Abstract
Recent work on continual learning in Large Language Models (LLMs) shows that on-policy finetuning is implicitly biased toward the solution of minimal KL divergence from the base model, allowing reverse-KL-driven objectives to learn new knowledge while largely preserving prior capabilities. We argue that unlearning can be viewed as the inverse analogue of continual learning: rather than adding a mode while preserving the rest of the distribution, it must remove a specific forget mode while leaving the rest of the distribution untouched. Existing unlearning methods, however, typically rely on off-policy objectives that push the model toward an idealized unlearning target, leading to utility collapse. Since KL divergence is itself the limit of the R\'enyi divergence family, and every positive-order R\'enyi divergence only attracts a model toward a target distribution, we introduce MARU (Mode-Avoiding R\'enyi Unlearning), a method that uses the negative-order R\'enyi divergence as a natural mode-avoiding objective for unlearning. Anchored at the pre-unlearning policy, this yields an objective whose geometry implicitly preserves model utility without requiring an explicit retain set. Extensive experiments on the TOFU and MUSE benchmarks show that MARU achieves state-of-the-art performance on the hardest forget split (10%) at 1B scale, while delivering up to better resistance to membership inference attacks than comparable distillation-based methods, though these gains are weaker on MUSE. These results suggest that reframing unlearning through divergence geometry, rather than heuristic gradient objectives, offers a principled and effective path toward safe, controllable and utility-preserving unlearning in LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.