Neuro-Symbolic Reasoning Framework for Long-Document Understanding
Abstract
Multimodal long-document understanding remains challenging because relevant evidence is sparsely distributed across lengthy and redundant documents. Existing retrieval-based methods optimize page-level relevance only, yet individually relevant pages may still fail to collectively provide all the evidence required by a question. Moreover, cross-page reasoning is often generated without explicit verification, allowing unsupported steps and logical gaps to remain hidden under incomplete evidence. We therefore formulate long-document understanding as a neuro-symbolic reasoning process that couples neural evidence discovery with symbolic verification. Neural models derive question-specific evidence requirements, localize complementary evidence across pages, and construct a structured reasoning hypothesis. Symbolic constraints evaluate evidence coverage, validate intermediate reasoning, and enforce answer conditions. Failed checks trigger a bounded repair loop to retrieve missing evidence and revise the reasoning hypothesis, while unresolved cases lead to abstention. Experiments on three multimodal long-document benchmarks demonstrate state-of-the-art performance, improved evidence localization and token efficiency, and interpretable evidence navigation with auditable verification traces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.