Length-Agnostic Sparse Attention for Unlimited Document Parsing
Abstract
Real-world applications such as financial analysis and contract review often involve documents spanning tens or hundreds of pages, making ultra-long document parsing increasingly important. Existing solutions face a fundamental trade-off between scalability and cross-page modeling, as page-by-page parsing loses cross-page dependencies while full-document modeling faces both prohibitive computation and an overwhelming visual context. Fixed-budget sparse attention provides a natural way to bound the context accessed by each query, but reliably selecting the few relevant blocks from an expanding multi-page context remains challenging. Crucially, we observe that document evidence exhibits a hierarchical sparsity pattern, with relevant evidence sparse across pages but locally concentrated within relevant pages. Leveraging this structure, we propose Length-Agnostic Sparse Attention (LASA), which enables reliable fixed-budget routing with an attention scope independent of document length, providing a scalable foundation for unlimited document parsing. LASA combines Page Canonicalization to exploit page structure, constrain the visual search space, and stabilize page-level positional modeling, with Token Bounding to reliably select a fixed budget of query-relevant visual blocks. Together, these designs enable efficient end-to-end cross-page modeling with stable performance as document length increases. We further introduce MarathonDocBench, a benchmark for hundred-page document parsing with substantially longer documents and richer cross-page structures than existing benchmarks. The resulting LASA-based parser consistently outperforms strong page-by-page pipelines and general-purpose large-scale VLMs on both MPDocBench and MarathonDocBench. By bounding the effective attention scope of each query, LASA keeps per-step decoding time approximately constant as document length increases and achieves up to speedup over full attention at the hundred-page scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.