VisExMEM: Visual Evidence-Based Explainable Multimodal Evaluation Metric
Abstract
Current image–text evaluation often relies on a global score, which gives limited insight into which semantic details are supported by the image. Explainable metrics provide rationales, but the visual evidence supporting those rationales can still be difficult to verify. We introduce VisExMEM, a reference-free evaluator that decomposes captions into atomic claims, localizes visual evidence, verifies each claim, and averages claim-level support into a global score. Explanations use the same evidence and decisions. We train Qwen3.5-based 4B and 9B models on 100K samples with multi-task SFT, counterfactual refinement, and evidence-aware GDPO. Across eight benchmarks, VisExMEM-9B achieves the highest descriptive composite among open-source evaluators in our comparison, exceeding VQAScore by 11.7 points while remaining competitive with GPT-5.6 Sol. We also introduce AGITA-Bench, containing 600 human-annotated image–text pairs and 8.7K atomic claims for global and claim-level evaluation. For untuned Qwen3.5 and GPT-5.6 Sol controls, decomposed prompting raises the same composite by 3.6–5.1 points relative to holistic prompting. Human evaluation shows that VisExMEM explanations outperform prior explainable metrics and approach GPT-5.6 Sol.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.