CoCoUN: Full-Trace Concept Unlearning in Multimodal Reasoning Models
Abstract
Machine unlearning in multimodal reasoning models aims to remove specific concepts while preserving general reasoning ability. However, existing methods focus primarily on final answers, overlooking target concepts disclosed in the preceding chain of thought. We reveal that this answer-centric paradigm leaves a systematic leakage path: once a target is generated as a discrete reasoning token, it becomes both visible output and autoregressive context for subsequent generation. Based on this insight, we propose cocoun, a novel full-trace concept-unlearning framework that removes a target before its discrete realization. cocoun replaces the target span with an adaptive continuous trajectory before its first token is emitted. Its joint objective learns the discrete-to-continuous transition, unlearns the target throughout the trajectory, and preserves subsequent reasoning through latent-exit continuation learning. Extensive experiments across multimodal reasoning backbones and diverse visual concepts demonstrate that cocoun substantially outperforms existing methods in full-trace unlearning while maintaining strong performance on unrelated tasks. Ablation studies and causal analyses further confirm the contributions of transition learning, continuous unlearning, and latent-exit continuation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.