Matryoshka Machine Unlearning in Classification and Generation
Abstract
Machine Unlearning (MU) removes the effect of data on a trained model without retraining from scratch. Earlier works operate on the full embedding width for unlearning. Therefore, nothing in them prevents deleted content from remaining decodable at a narrower feature scale. To address this, we propose Matryoshka Machine Unlearning (MMU), a nested prefix embedding that freezes the source model as a privileged multi-scale teacher. The student tries to mimic the teacher without having access to teacher-privileged information. The main idea is to arm the model with richer representations via coarse-to-fine nested embeddings. MMU is a plug-and-play method applicable on top of any MU method. Across three datasets, two backbones, three types of deletion, and fourteen unlearning baselines, MMU is the only method without a failure mode. Under class deletion, MMU drives forgetting accuracy to exactly 0% across all six dataset–backbone configurations while maintaining the highest retained accuracy among all baselines. Its real-world practicality has been demonstrated on the Physical Artificial Intelligence (PAI) dataset we collected. Additionally, for the image-generation task of removing dangerous content, it matches the best baseline, as shown in Fig. 1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.