CMMSafe: Rethinking MLLM Safety from Harmful Intent to Harmful Consequences
Abstract
While safety alignment for Multimodal Large Language Models (MLLMs) has gained significant attention, existing work addresses malicious intent and situational risks. We propose shifting the safety frontier toward consequence-driven safety, with a focus on latent hazards under benign requests. To formalize this shift, we introduce CMMSafe, a benchmark comprising 455 curated query-image pairs designed to evaluate a model's ability to identify latent hazards within context-dependent causal chains. Our analysis reveals frequent hazard omissions on this challenge set, with failure rates reaching 61.5% among the evaluated closed-source models, and limited or negative transfer under the evaluated static-alignment settings. To address these bottlenecks, we develop Consequence-Aware Safety Policy Optimization (CASPO), which combines constitution-conditioned token distillation from a fixed teacher with outcome rewards. Experimental results demonstrate that CASPO substantially improves risk appraisal, reducing the risk-identification failure ratio to 7.3% for Qwen2.5-VL-7B and 5.7% for Qwen3-VL-4B with task-dependent trade-offs in response effectiveness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.