acceptodds
Under review as a conference paper at ICLR 2027

Beyond Chunks and Binary Graphs: Multimodal Hypergraph Representation for Long Document Question Answering

Abstract

Long document question answering often requires grounding answers in textual, tabular, and visual evidence distributed across multiple pages, especially for visually rich documents such as scientific papers and technical reports. Existing Retrieval-Augmented Generation (RAG) approaches typically retrieve text chunks, multimodal chunks, full-page images, or binary graph structures, but these retrieval units provide limited support for organizing heterogeneous evidence into coherent retrieval units. We introduce DocHyper, a multimodal document hypergraph representation in which document entities are connected by hyperedges that encode n-ary evidence units within individual elements and among heterogeneous elements in the surrounding document context. Built from parsed document elements, DocHyper supports entity- and hyperedge-level retrieval with page- and layout-aware source grounding. Experiments on LongDocURL and MMLongBench-Doc show that our approach improves answer quality over chunk-, multimodal-chunk-, page-image-, and graph-based retrieval baselines. Ablations further analyze entity-level, hyperedge-level, and binary-graph retrieval, providing evidence for the value of organizing retrieved multimodal evidence as higher-order units with source-page grounding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.