acceptodds
Under review as a conference paper at ICLR 2027

StateHazard-MM: A Benchmark for Selective Unlearning of State-Dependent Hazards in Multimodal Models

Abstract

Multimodal large language models (MLLMs) deployed in scientific and clinical domains may encode hazardous visual knowledge, where the same scene can be benign or unsafe depending on subtle contextual states, making selective removal challenging due to strong visual overlap between hazardous and benign cases. Existing unlearning benchmarks focus on textual memorization or coarse visual concepts and do not capture this state-dependent setting. We introduce StateHazard-MM, a multimodal benchmark spanning six domains with a paired dual-split design that aligns hazardous samples for forgetting with semantically matched benign counterparts for retention. Built from over real-world images with LLM-assisted structured annotations, it enables scalable and controlled evaluation of state-sensitive unlearning. To address this challenge, we propose DynaRank-U, a rank-space unlearning framework that diagnoses each low-rank adaptation dimension using gradient amplitude ratio and directional conflict, and dynamically masks updates to suppress hazardous subspaces while preserving shared representations. Experiments across nine MLLMs show that DynaRank-U reduces hazardous response rates by 24%–35% while preserving over 95% of core reasoning capability (MMLU, MMMU-Pro) and substantially mitigating the dialogue degradation incurred by existing methods. Control experiments further verify that this forgetting is driven by state-sensitive semantics rather than stylistic cues, and that it is strongly grounded in the visual input. These results highlight the importance of structured, conflict-aware parameter selection for effective state-sensitive multimodal unlearning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.