acceptodds
Under review as a conference paper at ICLR 2027

RAMP: Residual-Guided Route Weighting for Vision-Language Unlearning

Abstract

Suppressing a target answer does not necessarily prevent a vision-language model from revealing the same fact through alternative wording or visual evidence. We introduce RAMP (Residual-Adaptive Multimodal Penalties), a two-stage unlearning method that uses residual answer scores to determine both when and where to apply further forgetting updates. Stage-I suppresses target answer spans. Stage-II scores verified alternative answers, activates representation penalties when those scores remain high, and allocates their strength across text, image, and fusion representations using gradient sensitivity. Alias suppression and preservation losses jointly constrain the updates. On CLEAR, Stage-II reduces held-out paraphrase recovery from 0.181 to 0.072 and indirect recovery from 0.174 to 0.089. On UMU-Bench, RAMP achieves lower remaining target performance and higher utility than the reported unlearning baselines at all three forget rates. Fixed-budget routing controls and baselines with shared supervision distinguish the contribution of adaptive weighting from that of additional data. Unseen attacks, output audits, and public-data substitution further characterize the method's effectiveness and practical requirements. Together, these results support using residual recovery to guide the allocation of unlearning updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.