Visual Activation-Guided Sparse Editing for Multimodal Large Language Model Unlearning
Abstract
Multimodal large language models (MLLMs) may memorize private, copyrighted, or unsafe visual information, making selective unlearning essential for their responsible deployment. Existing methods either update broad parameter regions or remove target-related neurons, which can damage representations shared with retained knowledge. To address this challenge, we propose Visual Activation-Guided Sparse Editing (VASE), a novel method for selective visual unlearning in MLLMs. VASE locates a small set of visual feed-forward network channels using normalized activation energy on forget images and learns sparse residual updates only for these channels. A bounded forgetting objective prevents excessive suppression, while output- and activation-level constraints preserve retained behavior. The learned updates are merged into the visual backbone without additional inference overhead. Experiments across multiple benchmarks and MLLM backbones show that VASE achieves state-of-the-art forgetting while preserving retained knowledge and general multimodal capabilities. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.