acceptodds
Under review as a conference paper at ICLR 2027

All the Facts, Still the Wrong Answer: Integration Gaps in Vision-Language Models

Abstract

Multimodal reasoning requires models to combine multiple pieces of evidence into decisions that satisfy all relevant constraints. Evaluations centered on final-answer accuracy leave unclear whether failures arise from reading individual facts or using them together. We study this question through the multi-evidence integration gap, linking local fact recovery to global decision validity on controlled subset-selection and path-finding tasks. Our evaluation combines conditional performance given correct local answers with same-grid local queries, text-oracle inputs, retrieval controls, and constraint ablations. We further examine hidden-state probes and activation replacement across network depth, token positions, and model families. Across the evaluated models, accurate local recovery coexists with substantial global failure. For example, Gemma-3-27B recovers all evaluated isolated item facts correctly but achieves only 19.75% global validity on eight-item visual instances. Probes decode information about solution membership from global-prompt states, while corresponding clean-state replacement recovers failed decisions over model-dependent depth ranges. These findings show that recovering facts and representing solution-related information are insufficient for reliable joint decisions. They motivate evaluations and model improvements that explicitly target how available evidence is combined, alongside advances in perception and model scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.