acceptodds
Under review as a conference paper at ICLR 2027

Retain-unlearn Gated Gradient Unlearning

Abstract

We propose TVE-Fisher, a simple and principled approach to unlearning in multimodal models that makes task-arithmetic editing selective. We cast unlearning as a retention-constrained local edit and observe an asymmetry of orders: forgetting grows linearly with the edit, whereas retention damage has no first-order term and is governed by Fisher information (FI). Retain curvature should therefore modulate the forget gradient. Under a diagonal geometry, this yields an optimal parameter-selection score, which we relax into an order-preserving soft mask with guarantees on optimality and robustness to estimation error. Concretely, TVE-Fisher computes retain FI statistics once, on a subset of observable retain data under the transductive protocol, to identify parameters critical to retained knowledge. It tracks forget FI statistics dynamically with an exponential moving average to capture the parameters currently relied upon for fitting the forget data. The resulting mask gates gradient updates during forget-side fine-tuning. We then extract a task vector from the difference between the initial and fine-tuned parameters and negate it; the same analysis shows that negation preserves the selected parameters and the retention cost while reversing the forget effect. Extensive experiments on standard vision–language model (VLM) and multimodal large language model (MLLM) unlearning benchmarks across multiple backbones show that TVE-Fisher achieves a stronger forgetting–retention trade-off than related baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.