Knowledge Regrows Across Modalities: Hidden Failure Mode in MLLM Unlearning
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable success in integrating information from various modalities, yet they raise growing concerns about privacy leakage. Machine unlearning is considered a remedy for privacy leakage, but existing multimodal unlearning approaches directly adopt techniques developed for unimodal models. This paper uncovers a hidden failure mode in MLLM unlearning, which we term *Cross-modal Knowledge Regrowth*: textual prompts can restore forgotten visual knowledge after unlearning, even in the absence of images. We identify that this phenomenon is attributed to the reactivation of vision-specific parameters by text-only prompts. Based on these findings, we show that selectively removing these vulnerable weights effectively suppresses cross-modal knowledge regrowth. Extensive experiments demonstrate that our approach reliably preserves unlearning robustness while maintaining model utility. Code is provided in the supplementary material for reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.