acceptodds
Under review as a conference paper at ICLR 2027

Don't Let Reasoning Slip Away: Mechanisms and Defense of Implicit Reasoning Risks

Abstract

Multimodal Large Language Models (MLLMs) face a distinct risk during long-chain reasoning, known as implicit reasoning risk, in which unimodal inputs that are benign in isolation combine to produce an unsafe output. We find that the recently emerging Thinking with Images (TwI) paradigm reduces this risk relative to the conventional text-chain-of-thought (Text-CoT) multimodal reasoning paradigm. To identify the cause of this behavior, we conduct an in-depth analysis and identify an important issue: the reasoning path of Text-CoT tends to move from concepts toward hypothetical or euphemistic formulations, which yields unsafe responses. We call this phenomenon Semantic Slippage Escape. TwI anchors the reasoning state to concepts and thereby mitigates this escape. Building on this insight, we propose AnchorGuard, a training-free adaptive defense framework that monitors, during early generation, the alignment between the hidden representation and an escape direction and converts this alignment into an intervention strength, which serves as the interpolation weight between the original representation and its projection onto the concept subspace. Experiments show that AnchorGuard improves the safety rate by 9.7% while preserving task utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.