acceptodds
Under review as a conference paper at ICLR 2027

NCAR: Neural Collapse Alignment Regularization for Machine Unlearning

Abstract

Machine unlearning aims to remove the influence of requested training data while efficiently approximating the solution induced by retraining on the retained data. Existing gradient-based methods primarily optimize predictive or output-level objectives for forgetting and retention, yet a discrepancy in the retained representation-classifier structure can remain even when predictive behavior is close to retraining outcomes. To address this structural discrepancy, we propose Neural Collapse Alignment Regularization (NCAR), which uses the idealized Neural Collapse terminal geometry of supervised classification as a structural prior. NCAR constructs a canonical simplex over the remaining classes, aligns it with the current classifier, and adaptively fits its translation and scale to retained features to obtain a classifier-aligned structural target. The resulting retain-side regularizer couples within-class concentration and simplex organization of class means with feature-classifier alignment. NCAR can be incorporated into existing gradient-based unlearning objectives while preserving the base method's forgetting-loss construction and parameter-update restrictions. We evaluate NCAR across random-subset and full-class unlearning, multiple datasets, architectures, and base methods, analyzing predictive behavior and retained representation structure together.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.