acceptodds
Under review as a conference paper at ICLR 2027

ReSight: Controlled Visual Reactivation after Multimodal Unlearning

Abstract

Clean-input evaluations of multimodal unlearning measure whether a target behavior is suppressed on an original image, but do not establish whether bounded image changes can elicit it. We study this gap as local target accessibility and introduce ReSight, a reference-conditioned visual reactivation attack. ReSight integrates reference-conditioned image construction, reference-free execution, and joint response controls. For each example, a fixed token prefix from a released response reference defines a negative-log-likelihood objective for projected sign-gradient updates under an L∞ bound. The question and model parameters remain unchanged. The resulting image is paired with the original question for free generation without the reference. Joint clean, unrelated-image, and blank-image controls identify target reactivation beyond clean-input persistence and transfer to other image bases. Across SafeEraser-240 and BeaverTails-V-108, two model families, and three forgetting recipes, ReSight recovers designated behaviors on 4.7–34.9% of clean-suppressed items, with control-qualified reactivation in all 12 settings covering 4.2–33.3% of full cohorts. Under the evaluated configurations, ReSight achieves the highest control-qualified success rate in 11 of 12 settings against five visual-attack adaptations. Fixed-image replay reveals checkpoint dependence. An exploratory spatial countermeasure reduces reactivation for fixed attacks. These findings demonstrate that clean-input suppression alone is insufficient evidence of robust multimodal unlearning and establish local target accessibility as a key dimension of its evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.