DocPilot: Interactive Agent Orchestration and Training for Long Visual Document Understanding
Abstract
Multimodal long-document question answering requires models to seek sparse evidence across pages and modalities, understand heterogeneous document content, and reason over the collected evidence. Even advanced vision-language models (VLMs) can overlook sparse evidence and struggle with cross-modal reasoning in a single pass. Existing document agents further introduce multi-step interaction, but some still rely on predefined workflows or coarse-grained actions, limiting the flexibility and granularity of document reasoning. We introduce DocPilot, an agentic document framework that not only supports expressive orchestration but also makes the complex document-solving process learnable by small language models (SLMs). DocPilot first orchestrates the following capabilities: evidence seeking through page-level navigation and fine-grained evidence localization; and multimodal understanding through local visual inspection, code-based numerical reasoning, and cross-modal reasoning. To transfer these capabilities to small models, we propose curriculum-based agentic training that progresses from single-ability learning to adaptive orchestration through supervision from data-grounded reasoning traces and executable interaction trajectories. Across MMLongBench-Doc, DocBench, FinRAGBench-V, and MADQA, DocPilot achieves strong gains with frontier VLM backbones, while DocPilot-4B delivers competitive performance against substantially larger models, demonstrating strong effectiveness and generalization across diverse long-document reasoning settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.