SURGE: Sinkhorn-Seeded Graph Expansion for Multimodal Long-Document Retrieval
Abstract
Retrieval-augmented generation over long documents is increasingly important in practice, yet most existing approaches are designed for purely textual corpora and fail to handle the heterogeneous mix of text, tables, and figures found in real-world technical documents. We present SURGE, a two-stage retrieval framework for multimodal long-context documents. In the first stage, Sinkhorn-regularized optimal transport is applied over all pairwise chunk similarities to produce a globally calibrated co-association graph that encodes inter-chunk relationships across modalities without privileging text. In the second stage, Personalized PageRank is seeded from a query-proportional personalization vector and propagates relevance through this graph, surfacing tables and figures that cosine-similarity retrieval systematically misses due to the embedding-space modality gap. We evaluate SURGE on long multimodal datasets spanning up to 620K tokens across vision-language models of different sizes and demonstrate advantages over baselines on document querying. Our analysis further identifies multimodal document segmentation as the critical open bottleneck: retrieval quality is fundamentally chunking-quality-dependent, and we call for research on robust multimodal chunking pipelines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.