CLEAR: Multi-Objective LLM Unlearning with Provable Performance Guarantees
Abstract
Large Language Model (LLM) unlearning has emerged as a crucial paradigm for mitigating privacy and security risks by aiming to erase undesirable knowledge while preserving the LLM’s general capability. Despite recent development of unlearning algorithms, existing methods suffer from critical limitations. First, they primarily focus on the basic trade-off between forgetting and retention, neglecting other essential dimensions such as robustness. Second, they overlook the failure-sensitive nature of unlearning at the task level, generally failing to provide worst-case performance guarantees. Third, current approaches are largely heuristic, lacking rigorous theoretical foundations and finite-time convergence analysis. In this paper, we systematically address the LLM unlearning problem through the lens of constrained multi-objective optimization (MOO), which enables us to propose Chebyshev-weighted LLM unlearning algorithm (CLEAR), a new penalty-based MOO algorithm with provable performance guarantees. Specifically, we theoretically construct a rigorous Karush-Kuhn-Tucker (KKT)-based metric, based on which we establish a finite-time convergence guarantee for the proposed CLEAR algorithm, marking the first of its kind in the LLM unlearning literature. Extensive numerical experiments across multiple benchmarks, employing comprehensive evaluation metrics, demonstrate that CLEAR significantly outperforms state-of-the-art baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.