acceptodds
Under review as a conference paper at ICLR 2027

MT-NPO: Modality-Aware Token-Level Negative Preference Optimization for MLLM Unlearning

Abstract

Multimodal large language model (MLLM) unlearning must suppress target knowledge while preserving retained utility, and this challenge is further complicated by uneven forgetting across modality-conditioned views. These difficulties expose two limitations that existing objectives do not explicitly address. At the token level, sequence-level optimization overlooks the unequal influence of individual answer tokens on future generation. At the modality level, a target may be suppressed under one view yet remain accessible through another. We identify these issues as sequence coupling and modality imbalance, and introduce MT-NPO, a modality-aware token-level unlearning framework that coordinates optimization at both levels. MT-NPO uses bounded FutureKL-guided suffix scores to focus unlearning pressure on influential tokens, while modality-specific unlearning and retention objectives adaptively allocate optimization according to residual target accessibility and utility degradation. Across multiple benchmarks and MLLM backbones, MT-NPO consistently improves the unlearning–retention trade-off and cross-modal balance over representative baselines, highlighting the importance of jointly coordinating unlearning across tokens and modality-conditioned views. Our code is available at https://anonymous.4open.science/r/MTNPO-DDD1/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.