acceptodds
Under review as a conference paper at ICLR 2027

ITJoint: Benchmarking and Diagnosing MLLMs on Query-Driven Image-Text Joint Extraction from Long Documents

Abstract

In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. Structuring such information into usable multimodal data is therefore a practical demand. Motivated by this demand, we systematically formulate a new research problem, query-driven image-text joint extraction from multimodal long documents, and develop a two-level taxonomy covering challenges arising from user intent and document content. Further, we construct ITJoint, a high-quality, manually annotated benchmark for systematically evaluating and diagnosing MLLM capabilities on this problem, comprising 2,455 pages of domain-specific documents with numerous non-decorative images and 910 answer instances. Our evaluation reveals three key limitations of current MLLMs: they tend to over-rely on visual similarity for image-text association, struggle to exhaustively recover all associated images in one-to-many correspondences, and exhibit a clear gap between page localization and fine-grained image localization. To examine whether these bottlenecks can be mitigated at the system level, we further introduce Q2IT as a strong multi-agent baseline. Our analyses show that multi-agent approaches can compensate for MLLM weaknesses and better leverage their strengths through specialized external tools and structured strategy design, while remaining vulnerable to error propagation across stages. The complete dataset, prompts, and code are publicly available at https://anonymous.4open.science/r/FullResources-72C7/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.