acceptodds
Under review as a conference paper at ICLR 2027

Blindfold-to-Forgetting: Route-Aware Multimodal Unlearning

Abstract

Multimodal unlearning aims to remove the influence of specific training data from models that process images and text while preserving performance on non-target data. Since the same target information can be elicited through distinct image- and text-based routes, intervention should account for their differences. Whereas prior work has primarily focused on editing modality-specific neurons or internal paths, we propose Blindfold-to-Forgetting, which uses the visual-to-language interface as a selective intervention point. The first stage shifts forget-image representations toward the mean representation of generic face images while preserving retained-image representations. The second stage suppresses target responses under image inputs processed by the original projector and text-only inputs, using a hinge objective and retain supervision to limit excessive suppression and utility degradation. Across three benchmarks and three MLLMs, our method achieves the highest Forget-Retain harmonic mean in 8 of 12 evaluation settings on UMU-Bench and CLEAR, with improvements of up to 18.1 points over the strongest baseline. Relearning experiments further show limited target-response recovery under the tested low-data conditions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.