M2E-X: Verified Evidence Graphs for Multi-Image Mathematical Reasoning
Abstract
Multi-image mathematical reasoning requires a model to identify entities across views, recover numerical and relational facts, and determine which visual evidence is reliable enough to support a solution. Existing multimodal reasoners typically operate directly on image-text representations, making it difficult to control evidence selection or diagnose when visual evidence is incomplete or contradictory. We introduce M2E-X, a two-stage framework that separates evidence construction from downstream reasoning. Stage 1 converts multiple images into an image-backed evidence graph containing grounded entities, OCR and value cues, intra-view relations, and cross-view correspondences. Stage 2 retrieves a sparse proof-oriented subgraph, aligns the selected evidence with the question, and verifies its completeness, semantic consistency, and recoverability before the evidence is exposed to a frozen vision-language reasoner. The framework further supports counterfactual verification and selective abstention when the available evidence cannot be safely released. This formulation turns multi-image mathematical reasoning into a controlled evidence selection and verification problem rather than unconstrained end-to-end generation. We evaluate M2E-X using both proof-level evidence metrics and downstream mathematical reasoning accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.