Reasoning-Aware Cross-Modal Association Unlearning in Multimodal Large Reasoning Models
Abstract
Multimodal large reasoning models (MLRMs) integrate visual inputs with textual and parametric knowledge, allowing sensitive information to be accessed through cross-modal associations during reasoning. This raises privacy concerns when a visual input should no longer provide access to such information. Machine unlearning offers a way to remove this influence without retraining the model from scratch. However, existing multimodal unlearning methods mainly target samples, concepts, or knowledge content, without explicitly separating the knowledge from the cross-modal association through which it is accessed. In reasoning models, this limitation is further exposed: target information may disappear from the final answer while remaining in intermediate reasoning, or simply be replaced by incorrect substitute knowledge. We therefore study cross-modal association unlearning, where the forget target is a designated visual-to-knowledge association: the model should suppress observable target-bearing behavior under the designated visual context while preserving access to the same knowledge in permitted contexts. We propose Visual-Keyed Counterfactual Rebinding (VKCR), a parameter-efficient framework for cross-modal association editing. VKCR first localizes the edit to a compact association-conditioned parameter subspace. It then reduces support for the designated target at task outputs and counterfactually rebinds the affected reasoning behavior by discouraging target-bearing or unsupported continuations while favoring valid non-sensitive alternatives. We further introduce CLEAR-R and MLLMU-R to evaluate target leakage, hallucinated reassociation, reasoning preservation, and modality-specific utility. Experiments show that VKCR substantially reduces image-conditioned target leakage and hallucinated reassociation, including across heterogeneous task formulations, while preserving non-sensitive visual reasoning and permitted text-only knowledge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.