MT-NPO: Modality-Aware Token-Level Negative Preference Optimization for MLLM Unlearning
Abstract
Multimodal large language model (MLLM) unlearning must suppress target knowledge while preserving retained utility, and this challenge is further complicated by uneven forgetting across modality-conditioned views. These difficulties expose two limitations that existing objectives do not explicitly address. At the token level, sequence-level optimization overlooks the unequal influence of individual answer tokens on future generation. At the modality level, a target may be suppressed under one view yet remain accessible through another. We identify these issues as sequence coupling and modality imbalance, and introduce MT-NPO, a modality-aware token-level unlearning framework that coordinates optimization at both levels. MT-NPO uses bounded FutureKL-guided suffix scores to focus unlearning pressure on influential tokens, while modality-specific unlearning and retention objectives adaptively allocate optimization according to residual target accessibility and utility degradation. Across multiple benchmarks and MLLM backbones, MT-NPO consistently improves the unlearning–retention trade-off and cross-modal balance over representative baselines, highlighting the importance of jointly coordinating unlearning across tokens and modality-conditioned views. Our code is available at https://anonymous.4open.science/r/MTNPO-DDD1/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.