acceptodds
Under review as a conference paper at ICLR 2027

Does Reasoning Make Multimodal Large Reasoning Models Safer?

Abstract

The widespread adoption of multimodal large language models (MLLMs) has also exposed safety vulnerabilities. Even relatively simple multimodal jailbreak attacks can induce them to generate harmful content. More recently, reasoning has been introduced to improve MLLMs' understanding and processing of multimodal tasks. However, recent studies report conflicting evidence about whether reasoning improves model safety. We find that reasoning does not always improve safety. It amplifies harmfulness under execution-dominant attacks but mitigates harmfulness under verification-dominant attacks. This contrast gives rise to a Reasoning Safety Reversal (RSR) in the safety effect of reasoning. We interpret this behavior through the Decode-Check-Execute (DCE) framework, which describes three coupled processes: recovering latent intent, checking safety, and executing the requested task. The safety outcome depends on which process dominates. We further design a Check-Before-Execute (CBE) prompting strategy to test this explanation. Based on these findings, we propose Intent-Gated Thinking Supervision (IGTS), a parameter-efficient LoRA-based method that explicitly supervises safety checking during reasoning. Evaluations across models and attacks show that IGTS improves the safety of multimodal large reasoning models (MLRMs).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.