DocArena: From Raw Documents to Search Training Data Beyond Answer Correctness
Abstract
Search agents need to locate and combine information from document collections through multi-step retrieval and reasoning. We study how existing documents can provide both answer targets and evidence labels for learning this capability. How- ever, correct question-answer pairs do not ensure accurate evidence labels, and incorrect labels can reward irrelevant retrieval or penalize the retrieval of support- ing evidence. To address this challenge, we propose DocArena, a fully automated pipeline that constructs questions, answers, and automatically checked evidence- page labels from raw documents without human training annotations or expert search demonstrations. Specifically, we extract textual and visual page content with a multimodal LLM and profile the distribution of facts across pages to guide evidence selection and multi-page question construction. The pipeline checks an- swer consistency and tests whether each question remains answerable after re- moving each evidence page in turn. We construct DocArena-79K with 79, 623 QA pairs from 8, 336 documents across 16 domains and 49 languages. We fur- ther develop Doc-Search, which decouples multimodal page perception from a text-based policy and uses the constructed answer and retrieval targets for rein- forcement learning. Experiments cover six multimodal document scenarios and seven text-based QA benchmarks. Pipeline ablations show that evidence checks improve training sample quality. Under fixed training settings, DocArena data improve document QA. Further analyses show that retrieval supervision improves evidence coverage and reduces repeated queries, while combining it with answer supervision improves QA. The trained policy also transfers from document col- lections to text-based QA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.