CoLU: Confidence-Weighted Layerwise Unlearning for Large Language Models
Abstract
Large language model (LLM) unlearning aims to remove targeted information while preserving useful capabilities without retraining from scratch. Methods such as NPO, SimNPO, and SatImp update the full model, while FOM-UL explores selective layer updates. Suppressing target information can damage retained knowledge, and this trade-off varies across layers even under the same unlearning loss. We propose Confidence-Weighted Layerwise Unlearning (CoLU), which combines probability-weighted forgetting with retain regularization while updating a single selected transformer layer. To reduce layer-selection cost, we introduce Brief Layer Screening (BLS), which ranks short unlearning trials on a subset of the training data using a normalized forgetting-retention (NFR) score. The shortlisted candidates then undergo independent full-data unlearning, and the highest-scoring model is selected. Compared with full screening, BLS reduces total training time by 60.2% to 79.7% on TOFU while preserving much of its performance. CoLU improves both forgetting and retention over NPO, SatImp, and GA-based FOM-UL in all six TOFU model and forget-split settings, and over SimNPO in five. On MUSE Books and News, CoLU achieves the highest retained KnowMem and the PrivLeak closest to zero among the compared unlearning methods. Across the measured TOFU and MUSE settings, average GPU memory usage for single-layer unlearning is 77.4% to 79.5% lower than for full-model unlearning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.