acceptodds
Under review as a conference paper at ICLR 2027

Tchebycheff Unlearn: Exploring the Pareto Front in Machine Unlearning

Abstract

Machine unlearning aims to forget target information while keeping the model useful. Applications differ in how much they need to forget and how much utility loss they can accept, necessitating a family of models with different trade-offs instead of a single model. We propose to generate the family by exploring the Pareto front in the forget-retain loss space, treating the forget loss and the retain loss as two objectives. A common approach is linear scalarization, which minimizes a weighted sum of the two losses. However, sweeping the weights cannot reach nonconvex regions of the Pareto front, so some trade-offs remain out of reach. In light of the above limitation of linear scalarization, recent work has also explored adaptive gradient methods from the multi-task learning literature, yet they offer no control over which Pareto-optimal solution they converge to. In this work, we introduce Tchebycheff Unlearn, which uses Tchebycheff scalarization to construct a family of models with different forget–retain trade-offs. We show that varying only the forget reference suffices to represent every attainable Pareto loss pair, with fixed positive task scales and retain reference. We evaluate our approach on the WMDP, TOFU, and MUSE datasets. Our empirical results show that Tchebycheff Unlearn improves both Pareto front coverage and the persistence of intermediate solutions compared with linear scalarization, enabling more comprehensive Pareto front exploration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.