acceptodds
Under review as a conference paper at ICLR 2027

Feature Reflection for Machine Unlearning without Retain Data

Abstract

Machine unlearning aims to remove designated knowledge while preserving a model's remaining capabilities. This preservation becomes more challenging when retain data are unavailable and unlearning relies solely on the pretrained model and the forget data. To overcome this challenge, we revisit unlearning from a Bayesian perspective and identify two requirements: reducing fit to the forget data while preserving pretrained knowledge. Motivated by these requirements, we introduce Feature Reflection (FR), a simple mechanism that turns bounded-below loss minimization into theoretically grounded forgetting. Specifically, FR reflects the features of the model being unlearned with respect to their pretrained counterparts, perturbs the reflection with multiplicative Gaussian noise, and minimizes the loss on the noisy reflected features. Theoretical analysis establishes that reducing the loss on reflected features raises a lower bound on the forget loss, explicitly connecting bounded-below minimization to forgetting. Further analysis shows that noise perturbation induces a curvature-weighted regularizer on deviations from pretrained features, promoting knowledge preservation. With a simple design and theoretical support, FR outperforms competitive retain-free unlearning baselines over eight vision datasets under different data settings, as well as on the WMDP benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.