acceptodds
Under review as a conference paper at ICLR 2027

Same State, Different Decisions: Disentangling Formation and Context-Conditioned Execution Mechanisms in VLMs

Abstract

Vision-language models (VLMs) often achieve higher accuracy on textual inputs than on semantically matched visual inputs describing the same task instance, with both modalities ultimately processed by the same decoder. We investigate this performance gap by causally delving into the decision process from modality-specific computation, through a cross-modal relay, to the downstream computation that continues to shape the final prediction. Our analysis identifies a relay state at which decision states transfer bidirectionally between visual and textual continuations. Yet identical relay states can still produce different predictions under different retained contexts, revealing that cross-modal executability can emerge before final decision commitment. This exposes a factorized decision process with distinct causal contributions from state formation and context-conditioned execution, while further localization attributes the residual execution effect primarily to post-relay attention reading distributed prefix information. Guided by these findings, we introduce Factorized Counterfactual Relay Distillation (FCRD), which distills Formation and Execution effects into a lightweight corrector that jointly updates the relay state and its surrounding context. Across three VLMs and five paired tasks, our method closes over 60% of the text-vision performance gap on average, demonstrating the consistency of the identified decision structure across model architectures and task types. Our analysis dissects the internal mechanisms underlying this gap and provides principled targets for correcting visual errors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.