acceptodds
Under review as a conference paper at ICLR 2027

RouteMask: Unlearning Identities Along Access Routes in VLMs

Abstract

Vision–language models (VLMs) learn from large collections of web images and text and are now widely used in assistants and search. Some of what they learn concerns real people: a single photograph can be enough for a model to reveal a person's birthplace, occupation, or education. Machine unlearning aims to remove such information after training. Existing methods for VLMs fine-tune the model with gradient ascent, preference optimization, or refusal targets, but they mostly change what the model says, not what it knows. After refusal training, the model still picks the true answer in multiple choice, and forgetting trained with names does not carry over to photographs. We address this in two steps. First, we locate where identity is accessed: injecting name states into frozen models while blocking attention to the image reveals an entity state that faces and names share in middle layers, and a visual bypass through which late layers read the image again. Second, we propose RouteMask, which updates only the MLP gate rows in these layers. At each step, a row mask compares the attribution of the current forget answers with fixed retain and perception profiles and scales the parameter update, and every forget term is paired with a retain term under both name and image cues. Experiments on MLLMU-Bench and S-MLLMUn with LLaVA-1.5 and Qwen2.5-VL show that both cues must be supervised and that forget and retain accuracy can only be separated when updates are restricted in this way.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.