acceptodds
Under review as a conference paper at ICLR 2027

Reasoning Models, Non-Reasoning Answers: Uncovering Reasoning-Bypass Shortcut in MLRMs

Abstract

Multimodal large reasoning models (MLRMs) aim to move vision language models beyond System-1-like direct prediction toward explicit System-2-like reasoning for complex problems. However, recent studies reveal substantial inconsistency between the reasoning process and the final answer. Existing work often characterizes such behavior using concepts borrowed from human cognition, including hallucination, unfaithfulness, and even deception, but provides limited insight into its internal mechanism. In this work, we identify a prevalent yet underexplored pattern of reasoning inconsistency, termed Reasoning-Bypass Shortcut, in which the final answer specifically reverts to the prediction the model would have produced without explicit reasoning. Across seven MLRMs and four benchmarks, Reasoning-Bypass Shortcut accounts for of all reasoning inconsistency cases, revealing a systematic tendency rather than random disagreement. To investigate its mechanism, we construct more than instances spanning diverse visual tasks that reliably capture this behavior. We then introduce a Reasoning Differential Probe to localize attention head representations associated with Reasoning-Bypass Shortcut, and further apply representation interventions to distinguish causal features from merely correlated ones. Our analysis reveals a sparse set of internal features that can redirect the model from its reasoning based decision (System-2) toward its direct prediction (System-1). Finally, we transfer our analysis to unseen scenarios and show that the identified internal signatures persist beyond the settings where they are first observed. This verifies the generality of our findings and shows that such behaviors can be captured from internal model representations, complementing output-level observations with a mechanistic perspective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.