acceptodds
Under review as a conference paper at ICLR 2027

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

Abstract

Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet assessing faithful use of textual and visual evidence in open-ended reports remains challenging. We introduce MMDeepResearch-Bench (MMDR-Bench), a benchmark of 140 researcher-curated tasks across 19 domains, where each task provides an image–text bundle to evaluate multimodal understanding and citation-grounded report generation. Compared to prior setups, MMDR-Bench emphasizes report-style synthesis with explicit evidence use, where models must connect visual artifacts to sourced claims and maintain consistency across narrative, citations, and visual references. We further propose a unified, interpretable evaluation pipeline: Formula–LLM Adaptive Evaluation (FLAE) for report quality, Trustworthy Retrieval-Aligned Citation Evaluation (TRACE) for citation-grounded evidence alignment, and Multimodal Support-Aligned Integrity Check (MOSAIC) for text–visual integrity, each producing fine-grained signals for aggregate failure-mode analysis beyond a single overall score. Experiments across 28 state-of-the-art models and agents reveal systematic trade-offs between generation quality, citation discipline, and multimodal grounding, highlighting that strong prose alone does not guarantee faithful evidence use and that multimodal integrity remains a key bottleneck for deep research agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.