FIT to Forget: Robust Continual Unlearning for Large Language Models
Abstract
While large language models (LLMs) exhibit remarkable capabilities, they increasingly face demands to remove the influence of privacy-sensitive, copyrighted, or harmful content. Existing unlearning methods primarily focus on single-shot scenarios, whereas real-world deletion requests arrive continually. Na\"ively applying these methods to sequential requests leads to severe utility degradation and catastrophic forgetting. To address this, we propose \fit, a robust framework for high-volume continual unlearning that resists both catastrophic forgetting and post-unlearning recovery. \fit stabilizes sequential updates through three synergistic mechanisms: redundancy Filtering, Importance-aware adaptive algorithm selection, and Targeted layer attribution. Furthermore, to facilitate rigorous evaluation, we introduce PCH, a unified benchmark spanning Personal information, Copyright, and Harmful content, alongside two symmetric metrics, Forget Degree (F.D.) and Retain Utility (R.U.), that quantify forgetting-utility trade-offs. Experiments across five LLMs (up to 14B parameters) show that \fit maintains a strong forgetting–utility balance throughout long request streams. Even after hundreds of randomly ordered requests, \fit preserves competitive downstream (\eg, GSM8K and MMLU) performance and resists relearning and quantization recovery attacks.Our code is available at https://anonymous.4open.science/r/FIT_Code-7304/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.