Evaluating Reasoning under Multimodal Conflict: Separating Score Gains from Evidence Selection
Abstract
Reasoning is used to help multimodal large language models answer questions when visual and textual evidence conflict. Yet higher benchmark scores do not establish more reliable evidence selection, because scores also depend on how answers are elicited and judged. We develop a controlled evaluation protocol that examines correct-source direction, candidate access, and answer scoring through complementary paired comparisons. Its core design crosses reasoning on/off, repeated versus disagreeing incorrect statements, and candidate visibility while holding the required response and output limit fixed within each model. The core comparison covers three open-weight model families, with separate evidence-direction and visual-recognition controls on two. We find that reasoning can improve answers when images are correct while harming them when text is correct. On image-associated knowledge questions, candidates enlarge the reasoning-by-disagreement interaction mainly through a weaker non-reasoning baseline in Qwen, but more through reasoning-enabled accuracy changes in Gemma. The large positive candidate-induced increases in this interaction observed in these two models are not reproduced in the third or under the chart-reading protocol. In the recognition control, the original strict-JSON tests fail, while supplementary wrapper-tolerant scoring reveals label-list benefits and no additional accuracy gain from constrained decoding. These results distinguish format compliance from label matching without attributing vocabulary gains to naming alone. Our findings support evaluating reasoning through source-resolved outcomes, absolute paired accuracies, and explicit answer-information and scoring controls, rather than aggregate gains alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.