acceptodds
Under review as a conference paper at ICLR 2027

CoCoUN: Full-Trace Concept Unlearning in Multimodal Reasoning Models

Abstract

Machine unlearning in multimodal reasoning models aims to remove specific concepts while preserving general reasoning ability. However, existing methods focus primarily on final answers, overlooking target concepts disclosed in the preceding chain of thought. We reveal that this answer-centric paradigm leaves a systematic leakage path: once a target is generated as a discrete reasoning token, it becomes both visible output and autoregressive context for subsequent generation. Based on this insight, we propose cocoun, a novel full-trace concept-unlearning framework that removes a target before its discrete realization. cocoun replaces the target span with an adaptive continuous trajectory before its first token is emitted. Its joint objective learns the discrete-to-continuous transition, unlearns the target throughout the trajectory, and preserves subsequent reasoning through latent-exit continuation learning. Extensive experiments across multimodal reasoning backbones and diverse visual concepts demonstrate that cocoun substantially outperforms existing methods in full-trace unlearning while maintaining strong performance on unrelated tasks. Ablation studies and causal analyses further confirm the contributions of transition learning, continuous unlearning, and latent-exit continuation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.