Does Reasoning Make Multimodal Large Reasoning Models Safer?
Abstract
The widespread adoption of multimodal large language models (MLLMs) has also exposed safety vulnerabilities. Even relatively simple multimodal jailbreak attacks can induce them to generate harmful content. More recently, reasoning has been introduced to improve MLLMs' understanding and processing of multimodal tasks. However, recent studies report conflicting evidence about whether reasoning improves model safety. We find that reasoning does not always improve safety. It amplifies harmfulness under execution-dominant attacks but mitigates harmfulness under verification-dominant attacks. This contrast gives rise to a Reasoning Safety Reversal (RSR) in the safety effect of reasoning. We interpret this behavior through the Decode-Check-Execute (DCE) framework, which describes three coupled processes: recovering latent intent, checking safety, and executing the requested task. The safety outcome depends on which process dominates. We further design a Check-Before-Execute (CBE) prompting strategy to test this explanation. Based on these findings, we propose Intent-Gated Thinking Supervision (IGTS), a parameter-efficient LoRA-based method that explicitly supervises safety checking during reasoning. Evaluations across models and attacks show that IGTS improves the safety of multimodal large reasoning models (MLRMs).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.