Can MLLMs Retain More Without Unlearning Less?
Abstract
Multimodal large language model (MLLM) unlearning aims to make MLLMs for- get chosen facts, such as a biography recalled from a portrait or a name, while retaining the rest of their knowledge. Existing benchmarks take a static, single- point view, scoring each method once at the end of its training. Even a fully converged run marks one position on the trade-off between forgetting and reten- tion and cannot characterize the method as a whole. We instead take a dynamic view and introduce iso-forget evaluation, which characterizes each method by its forget–retain frontier, the best retention it attains at every forgetting level. Unlike a single score, the frontier compares methods at the same forgetting level. It sep- arates a genuine improvement, which shifts the frontier, from a gain that merely trades forgetting for retention along it. Specifically, we start from a few anchor checkpoints at different unlearning strengths and reach any target forgetting level exactly, without retraining, by randomizing between two adjacent anchors with a closed-form probability. Across six MLLMs from three families, forgetting costs almost no retention up to a knee, whereas beyond it retention drops sharply. Tun- ing the training procedure or switching the objective mostly moves a model along the frontier, so the algorithm chiefly picks a point on it. The frontier itself is set by how the model represents and retrieves the forgotten facts, since forgetting spills over onto look-alike people and blocks mainly the cue it was trained on. MLLMs can therefore retain more without unlearning less by reaching a fact through every cue that recalls it, rather than by pushing harder on one.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.