PageIndex-Guided Multimodal GraphRAG: A Critical Evaluation of Structure-Aware Retrieval on Visually Rich Financial Documents
Abstract
Financial annual reports pack critical information into tables and charts, yet most retrieval-augmented generation (RAG) systems, graph based or not, only read the surrounding text. We evaluate whether adding a knowledge graph and multimodal (text plus vision) processing improves retrieval on real, visually rich documents, or whether it primarily adds complexity. We build PIM-GraphRAG, a system combining page-level hierarchical navigation with a schema-constrained multimodal knowledge graph and training-free cross-modal entity linking, and evaluate it against four baselines on ten real UK annual reports. The proposed system underperforms simple BM25 keyword search by 35 points on document accuracy (58.82% vs. 94.12%), and cross-modal linking recovers only 16 of 2,023 candidate entity pairs. A companion ablation across text-only, multimodal, and graph-augmented baselines shows why an unconstrained relation schema collapses distinct facts into one generic label, driving retrieval precision toward zero. We present these negative results as the paper's primary contribution, since they identify concretely where multimodal graph augmentation helps and where it does not, and we release the full pipeline for reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.