acceptodds
Under review as a conference paper at ICLR 2027

Models See More Than They Say: The Answer Bottleneck in Multimodal Language Models

Abstract

Multimodal language models can encode useful visual information and still answer incorrectly. We investigate this gap through the Guided Answer Space (GAS), an explicit set of candidate answers specifying the alternatives the model must distinguish. By varying this space while keeping the image, question, and target answer fixed, and intervening on intermediate representations, we provide evidence for an Answer Bottleneck in determining correct answers from encoded visual evidence. Crucially, intermediate visual representations already provide substantial support for accurate open-ended answers: they help recover accuracy lost to image corruption and retain much of their contribution even when their states are held fixed in later layers. Across eleven models on TextVQA and GQA, two-choice selection depends on visual tokens through fewer layers than open-ended answering, with a median difference equivalent to 42% of model depth; expanding the GAS extends this dependence to later layers. Building on these findings, we develop a two-stage inference procedure that constructs a GAS from other models' answers and then selects the final answer from it, without ground-truth access during either stage. Using three of the models, this procedure corrects 23.7% of original open-ended errors and improves overall accuracy by 3.56 percentage points. Our findings show that answer requirements shape the decoder's reliance on visual evidence and that the GAS can improve accuracy through two-stage inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.