acceptodds
Under review as a conference paper at ICLR 2027

Which Computations Drive Visual-to-Text Transfer? A Causal Ledger for Vision-Language Models

Abstract

Vision-language models can incorporate visual evidence into text states, but how much each decoder block contributes to the answer-relevant visual effect remains difficult to quantify. We introduce reciprocal nested interventions that decompose a common image-conditioned answer effect into ordered block contributions. By replacing visual rows at successive decoder boundaries while retaining text states, we measure how much of the visual effect has entered text states by each depth. These contributions sum exactly to the total effect, yielding a layerwise ledger of cumulative transfer and residual visual dependence. On Qwen3.5-0.8B and Qwen3.5-4B, just two dominant writer blocks carry and of the total visual effect, respectively. This rapid accumulation establishes early causal closures at (0.8B) and (4B). The ledger directly guides compression: closure determines when visual tokens can be evicted, while writer share prioritizes which blocks' inputs to protect in BF16 under quantization. These predictions depend on task and readout: transfer schedules shift across tasks, and generation can still require late visual access beyond an immediate-answer closure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.