acceptodds
Under review as a conference paper at ICLR 2027

CMMSafe: Rethinking MLLM Safety from Harmful Intent to Harmful Consequences

Abstract

While safety alignment for Multimodal Large Language Models (MLLMs) has gained significant attention, existing work addresses malicious intent and situational risks. We propose shifting the safety frontier toward consequence-driven safety, with a focus on latent hazards under benign requests. To formalize this shift, we introduce CMMSafe, a benchmark comprising 455 curated query-image pairs designed to evaluate a model's ability to identify latent hazards within context-dependent causal chains. Our analysis reveals frequent hazard omissions on this challenge set, with failure rates reaching 61.5% among the evaluated closed-source models, and limited or negative transfer under the evaluated static-alignment settings. To address these bottlenecks, we develop Consequence-Aware Safety Policy Optimization (CASPO), which combines constitution-conditioned token distillation from a fixed teacher with outcome rewards. Experimental results demonstrate that CASPO substantially improves risk appraisal, reducing the risk-identification failure ratio to 7.3% for Qwen2.5-VL-7B and 5.7% for Qwen3-VL-4B with task-dependent trade-offs in response effectiveness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.