Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
Abstract
Machine unlearning enables pre-trained models to forget specific data in compliance with privacy regulations. While approximate unlearning methods work well for class-wise forgetting, random data forgetting, which involves removing individual samples while retaining others in the same class, remains highly challenging, especially for Vision Transformers (ViTs). Through selective masking experiments, we discover that masking a small proportion of the highest-attention patches preserves ViT's recognition capability while significantly degrading its memorization. Building on this insight, we propose Lethe, a contrastive unlearning method tailored for ViTs that uses masked images as positive logits and original images as negative logits, guiding the model to forget sample-specific details while retaining category-level outlines. Experimental results on multiple benchmarks and ViT architectures show that Lethe surpasses existing approximate unlearning methods, achieving the closest gap to the retain-only Retrain reference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.