acceptodds
Under review as a conference paper at ICLR 2027

Causal Readout Dependence in Multimodal Language Models

Abstract

A multimodal model can identify an object when choosing among named alternatives yet fail to mention the same object when asked to describe an image. We ask whether this difference arises only from competition among possible outputs, or whether image-conditioned internal states themselves have different causal effects across readouts. Using Qwen, Mistral, and Llama vision–language models, we identify hidden states at text-token positions that change substantially with the image while preserving the concept named in text. We then transplant these image-conditioned states into otherwise image-absent runs and measure their effects across readouts. Under forced choice, late-layer transplantation recovers 89.1% of the image effect in Qwen, 82.8% in Mistral, and 48.0% in Llama. At the next-token position of a generation prompt, the same intervention recovers only 14.8%, 0.3%, and 7.4% of the image-induced increase in probability for the depicted object. Transplanting states across layers 0–31 increases recovery but leaves a large gap, and the depicted object is detected in none of 464 neutral transplanted captions per model family. Controls show that this difference is not explained by answer-space size, candidate naming, scoring procedure, or transplant coverage. Conversely, removing continued access to the image sharply disrupts generation while preserving much of the relative structure of scored decisions. We call this causal readout dependence: an internal intervention can have a large effect under one readout and a much smaller effect under another. The gap is already present before the first token is generated, and competition during generation widens it. Causal effects established under one multimodal readout therefore should not be assumed to transfer to another.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.