acceptodds
Under review as a conference paper at ICLR 2027

Do VLMs Fail at Causal Reasoning or Output Format Production?

Abstract

Evaluations of vision-language model (VLM) causal reasoning typically rely on a single output format, making it difficult to distinguish failures of format production from failures of reasoning. We introduce a diagnostic framework that combines structure-first generation, multi-format testing across four structural representations, three levels of prompt scaffolding, and score-compliance pairing, where content quality is reported alongside format compliance. Across six VLMs spanning 2023–2025 architectures, we find that the structural expression of causal reasoning is highly sensitive to format and prompt support. The same model reaches 97% compliance for numbered sequences but below 1% for arrows, while removing scaffolding yields compliance drops of 25–91 points. For five of the six models, structural and textual quality converge within the same response once the format barrier is removed. Targeted fine-tuning further helps distinguish limitations arising from insufficient training signal from those that persist under the tested architecture.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.