MM-ReportBench: A Comprehensive Benchmark for Multimodal Report Generation from Tables
Abstract
The ability to generate multimodal reports integrating fluent narratives with faithfulvisualizations from tabular data, has become a cornerstone capability for mod-ern vision-language models in practical domains such as finance, healthcare, andpolicy making. However, existing benchmarks are limited by single-modalityfocus, simplistic metrics, novelty evaluation, and visualization quality. We ad-dress these limitations with MM-ReportBench, a multimodal report evaluationbenchmark covering 386 tasks across 185 real-world tables. Each task is pairedwith an expert-curated reference report featuring standardized chapter structuresand verifiable key insight points, constructed through a hybrid pipeline of multi-modal reports generation based on Monte Carlo Tree Search and rigorous humanquality control. For robust evaluation, we design a dual-judge evaluation frame-work: Text Judge Model assesses textual quality along structural completeness,grounded insight verification with numerical fact-checking against source tables,and novelty analysis that distinguishes genuine analytical discovery from triv-ial paraphrasing; Visualization Judge Model evaluates generated visualizationsacross data fidelity, expressiveness, aesthetics, and technical correctness. They alsoevaluate cross-modal consistency. Extensive experiments on 20 state-of-the-artmodels reveal critical weaknesses: even the strongest model achieves only 68.9overall score, with numerical hallucination affecting 38.5% of reports and gen-uinely novel insights remaining rare. Our ablation studies further demonstratethat single-judge evaluation consistently inflates scores by 8–12 points comparedto dual-judge assessment, confirming that rigorous and reliable benchmarkingdemands multi-faceted, complementary evaluation probes. The source data and code is available at https://anonymous.4open.science/r/MReportBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.