RECAST: Beyond Weight Reinitialization for Machine Unlearning
Abstract
Machine unlearning is commonly evaluated through the behavior a model exhibits after data is removed. Yet similar behavioral outcomes can be achieved through completely different model modifications. This raises a question: what kind of internal state does an unlearning procedure actually produce? We study this question across classification and generative concept removal, comparing unlearned models on their outputs, their internal state, and how quickly forgotten behavior returns. Across two class-level classification benchmarks, methods with similar forgetting at the output reach substantially different internal states, often staying close to the original model, and differ in how fast the forgotten behavior comes back. Motivated by these differences, we introduce RECAST (REtain-guided Contraction via Alignment Scaled Transformation), a retain-guided weight intervention that uses gradient–weight alignment to continuously contract structural weight groups. RECAST combines strong behavioral agreement with the retrained reference and slower recovery of forgotten behavior than most evaluated baselines. Under repeated deletion requests, it is the only evaluated method whose features end close to a model retrained on the remaining classes. We further extend RECAST to nudity concept removal in FLUX, where it strongly suppresses the concept under both standard and adversarial prompts while preserving benign generation quality. Together, our results show that behavioral outcomes, internal state, and recoverability capture complementary properties of unlearned models, and position retain-guided continuous contraction as an alternative to reset-based weight interventions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.