acceptodds
Under review as a conference paper at ICLR 2027

MIRRA: Identity-Routed Refusal Adaptation for Unlearning in MLLMs

Abstract

Multimodal large language models (MLLMs) can memorize personal information and reveal it through text and image queries, motivating multimodal machine unlearning to protect individual privacy. However, existing methods that modify shared model parameters can degrade performance on unrelated queries, while refusal training alone does not ensure that target individuals are consistently recognized across modalities. To address these limitations, we propose MIRRA, a multimodal unlearning framework that separates identity recognition from refusal training. The key idea is to use identity recognition to determine when to activate a lightweight refusal adapter. MIRRA combines textual identity matching with face verification using a single reference image per identity, drawing on both forget and retain identity prototypes to distinguish target individuals from similar non-targets. A match in either modality activates the adapter to refuse the query, while other queries are handled directly by the base model. Experiments on MLLMU-Bench and CLEAR across three MLLMs show that MIRRA effectively suppresses responses about target individuals while preserving model utility. On MLLMU-Bench, MIRRA reduces forget-set classification accuracy to 0.42–1.25%, with retain-set classification accuracy decreasing by at most 1.23 percentage points relative to the original models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.