DocSEAM: Bidirectional Cross-Page Encoding for Multi-Page Document Understanding
Abstract
Multi-page document vision-language models encode each page independently, creating a feature-level discontinuity at page boundaries: visual tokens are formed without access to content on adjacent pages, and the language model must reason over these page representations formed in isolation. We introduce DocSEAM (Document Selective Enrichment via Asymmetric Masking), a lightweight module inserted between a pretrained vision encoder and the language model that injects the inductive bias of local cross-page context into visual representations. For each page, DocSEAM forms a local triplet with its two neighbors, applies asymmetric masked self-attention with directional salience gates to fuse information from neighbor pages, and returns only the enriched center-page tokens, leaving the visual token count unchanged. After a supervised cross-page warm-up and multi-page fine-tuning with DocSEAM frozen, our 9B-parameter model achieves state-of-the-art results among open-weight models on five multi-page document VQA benchmarks. We further show that DocSEAM makes the model robust to arbitrary re-pagination and significantly improves cross-page span localization, confirming that the gains stem from improved cross-page representations
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.