acceptodds
Under review as a conference paper at ICLR 2027

CoLU: Confidence-Weighted Layerwise Unlearning for Large Language Models

Abstract

Large language model (LLM) unlearning aims to remove targeted information while preserving useful capabilities without retraining from scratch. Methods such as NPO, SimNPO, and SatImp update the full model, while FOM-UL explores selective layer updates. Suppressing target information can damage retained knowledge, and this trade-off varies across layers even under the same unlearning loss. We propose Confidence-Weighted Layerwise Unlearning (CoLU), which combines probability-weighted forgetting with retain regularization while updating a single selected transformer layer. To reduce layer-selection cost, we introduce Brief Layer Screening (BLS), which ranks short unlearning trials on a subset of the training data using a normalized forgetting-retention (NFR) score. The shortlisted candidates then undergo independent full-data unlearning, and the highest-scoring model is selected. Compared with full screening, BLS reduces total training time by 60.2% to 79.7% on TOFU while preserving much of its performance. CoLU improves both forgetting and retention over NPO, SatImp, and GA-based FOM-UL in all six TOFU model and forget-split settings, and over SimNPO in five. On MUSE Books and News, CoLU achieves the highest retained KnowMem and the PrivLeak closest to zero among the compared unlearning methods. Across the measured TOFU and MUSE settings, average GPU memory usage for single-layer unlearning is 77.4% to 79.5% lower than for full-model unlearning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.