acceptodds
Under review as a conference paper at ICLR 2027

MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA

Abstract

Long-document multimodal question answering requires evidence whose amount, location, and modality vary with the question. One-shot document RAG typically retrieves a fixed number of text chunks or page images, leaving limited scope to recover missing clues, inspect local details, or remove irrelevant content before answering. Constructing sufficient evidence within a finite budget therefore requires query-dependent decisions about both the extent and granularity of the reader context. We propose MAGE-RAG for budgeted query-time evidence subgraph construction. It connects pages and local elements in an offline graph, uses an online controller to select and revise evidence from retrieved entry pages, and renders the resulting state as page-organized multimodal input. MAGE-RAG achieves 52.75 accuracy on LongDocURL and 53.26 accuracy with 51.19 F1 on MMLongBench-Doc. Breakdown and ablation results support the value of organizing cross-page and heterogeneous evidence; budget experiments further show that increasing the initial page count or controller budget does not consistently improve accuracy. Our code is available at https://anonymous.4open.science/r/MAGE-RAG-F8B3.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.