Separability Determines Feasibility, but Not the Rank: The Ceiling of Parameter-Space Unlearning and the Unpredictability of Its Rank Hyperparameter
Abstract
SEMU-style SVD-truncation unlearning hits a ceiling on CNN classifiers: over the whole range of the rank hyperparameter on CIFAR-100 whole-class forgetting the selected rank stays at 1 and forget accuracy remains pinned at 98.00%, and neither full rank nor extending to all layers restores forgetting. The geometric root of this failure is a class-mean common factor shared by the decision-layer gradients: under the unified convention , forget and retain inherit the same —on the random-subset experiment the two sides' mean features differ by a measured —leaving rank-1 updates no selective foothold. At the same depth, activations become separable once globally centered: their AUC margin over an in-distribution null ( to for whole classes, for random subsets) yields—calibrated per configuration, with no cross-configuration threshold an ex ante feasibility criterion—whose verdicts stay consistent on ImageNet and ViT experiments. We further prove that the projection rank has no static predictor: with activations, weights, and projection basis fixed, changing only the fine-tuning process moves from 4 to beyond 10—rank selection is necessarily a deployment-time scan.the criterion's practical value lies precisely in confining this scan cost to feasible problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.