acceptodds
Under review as a conference paper at ICLR 2027

EVIDENCE-GUIDED MULTIMODAL LARGE LANGUAGE MODELS VIA VISUAL CAUSAL REASONING

Abstract

Multimodal large language models (MLLMs) can answer open-ended questions about images, yet their outputs can be driven by linguistic priors that are only weakly supported by the visual input. We introduce EG-MLLM, an evidence guided architecture that inserts an explicit, testable evidence path between the vision encoder and language decoder. A prompt-conditioned selector identifies task-relevant visual tokens; a sparse structured reasoner models their dependen cies; counterfactual token interventions distinguish answer-critical evidence from matched nuisance context; and a bottleneck cross-attention adapter injects the resulting representation into the decoder. The training objective combines autore gressive likelihood, evidence consistency, and an intervention margin that requires removing selected evidence to affect the target likelihood more than removing nuisance tokens. We evaluate the implemented evidence pathway with a controlled 6,000-instance evidence-binding experiment and four 2,000-instance test conditions defined by fixed shortcut and noise settings. EG-MLLM reaches 98.65% accuracy under a deliberately reversed shortcut correlation and 95.40% under heavy nuisance corruption while maintaining 98.33% and 94.80% evidence precision, respectively. A matched run without the intervention objective obtains nearly identical answer accuracy but satisfies the intervention margin on only 49.65% of shortcut-shift cases, compared with 98.30% for the full model. These results show that sparse evidence selection, structured relation binding, and intervention-sensitive training can jointly improve robustness and make the model’s supporting evidence explicitly testable

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.