From Structured Facts to Logical Conflicts Visual-Text Conflict
Abstract
Large multimodal models are typically evaluated on image-text inputs that are semantically consistent, leaving their ability to handle visual-textual contradictions underexplored. We introduce LogicCon, a benchmark for visual-textual logical conflict reasoning that requires models to detect conflicts, localize the conflicting fact pair, and diagnose the conflict type. LogicCon constructs controllable contradiction samples by converting visually grounded question-answer pairs into consistent declarative statements, applying structured fact mutations, and verifying each mutation against symbolic visual evidence. We evaluate 13 representative open-source and closed-source multimodal models under a unified zero-shot protocol. Results show that current models can often detect whether a conflict exists, but perform substantially worse when required to recover the complete conflicting fact pair or reason over compositional samples. A large gap between the strongest models and human performance further indicates that visual-textual logical conflict reasoning remains a challenging and unresolved capability. Comparative experiments further reveal that hallucination mitigation methods mainly improve basic conflict detection but provide limited gains in recovering complete conflicting fact pairs, while general reasoning methods still struggle with precise visual evidence grounding. By explicitly parsing target-aware claims, verifying visual evidence, and aggregating conflicting facts, our method achieves more consistent improvements on higher-level logical conflict reasoning tasks. Code and benchmark artifacts are available at https://anonymous.4open.science/r/LogicCon-ICLR2027/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.