acceptodds
Under review as a conference paper at ICLR 2027

CUSE: Fine-Grained Multimodal Unlearning with Counterfactual Guidance and On-Policy Self-Distillation

Abstract

Multimodal large language models (MLLMs) can reproduce sensitive, copyrighted, or unsafe information from their training data. Effective unlearning must remove target information while preserving non-target content, even within the same response. Existing methods typically apply uniform forgetting objectives to complete responses and rely on fixed, off-policy supervision, resulting in coarse token control and misalignment with the model’s current behavior. We propose CUSE, a Counterfactual-guided Unlearning framework via on-policy SElf-distillation for multimodal large language models. It first constructs counterfactual inputs by removing target-related evidence and uses the resulting changes in token probabilities to identify which tokens require stronger forgetting. It then distills token-level forgetting and retention targets on responses sampled from the current policy, with a frozen copy of the original model serving as the teacher model augmented with the corresponding forgetting- and retention-oriented privileged contexts. This design jointly determines what to forget and what to preserve. On MLLMU-Bench, CUSE reduces forget-set ROUGE-L by 33.6% while achieving 1.2% higher retain-set ROUGE-L than the strongest baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.