acceptodds
Under review as a conference paper at ICLR 2027

CAVE-SMR: Contextualized Visual Evidence for Safety-Monotonic Multimodal Situational Reasoning

Abstract

Images and instructions that appear innocuous in isolation can lead to unsafe real-world consequences when combined, creating implicit cross-modal risks that existing vision-language models (VLMs) often overlook. Existing approaches struggle to establish a reliable connection between decisive visual conditions, the user's intended action, and its real-world consequences. Moreover, stronger safety interventions can induce over-refusal, creating an undesirable trade-off between safety and utility. Multimodal situational safety requires not only identifying risk-relevant visual evidence, but also connecting that evidence to the user's intended action and its real-world consequences. Based on this insight, this paper introduces a multimodal situational-safety framework called Contextualized Adoption of Visual Evidence with Safety-Monotonic Responses (CAVE-SMR). It first extracts request-relevant visual facts without presupposing a safety label and grounds them in a structured scene event that links the user's intended action, the affected object, and the resulting real-world consequence. A shared lightweight relation scorer then integrates the structured evidence produced by the preceding stages to estimate the risk of the complete event and compares it with a counterfactual event in which the decisive visual condition is removed, measuring whether that condition changes the safety judgment. For events identified as risky, the proposed method compiles the selected evidence into a structured risk record and applies risk-grounded generation followed by a deterministic continuation check, preventing the final response from reintroducing the hazardous objective while still providing concrete, actionable safe alternatives. Systematic experiments across three VLMs and three multimodal safety benchmarks show that CAVE-SMR consistently improves the recognition of implicit cross-modal risks and safe response behavior while preserving useful interactions on benign requests. It outperforms strong safety-enhancement baselines and achieves a more favorable balance between safety and utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.