Beyond Chunks and Binary Graphs: Multimodal Hypergraph Representation for Long Document Question Answering
Abstract
Long document question answering often requires grounding answers in textual, tabular, and visual evidence distributed across multiple pages, especially for visually rich documents such as scientific papers and technical reports. Existing Retrieval-Augmented Generation (RAG) approaches typically retrieve text chunks, multimodal chunks, full-page images, or binary graph structures, but these retrieval units provide limited support for organizing heterogeneous evidence into coherent retrieval units. We introduce DocHyper, a multimodal document hypergraph representation in which document entities are connected by hyperedges that encode n-ary evidence units within individual elements and among heterogeneous elements in the surrounding document context. Built from parsed document elements, DocHyper supports entity- and hyperedge-level retrieval with page- and layout-aware source grounding. Experiments on LongDocURL and MMLongBench-Doc show that our approach improves answer quality over chunk-, multimodal-chunk-, page-image-, and graph-based retrieval baselines. Ablations further analyze entity-level, hyperedge-level, and binary-graph retrieval, providing evidence for the value of organizing retrieved multimodal evidence as higher-order units with source-page grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.