DocComp: Query-Conditioned Semantic Component Selection for Efficient Document VQA
Abstract
Document visual question answering (DocVQA) requires high-resolution inputs to read small text and structured content, but full-page processing turns even question-irrelevant regions into visual tokens. Existing visual-token pruning reduces the sequence only after visual processing, whereas aggressive region selection can lose answer-supporting context, and both can be tied to a specific VLM. To address this, we propose DocComp, a plug-and-play framework for query-conditioned semantic component selection that keeps only question-relevant layout regions before downstream VLM processing. DocComp decomposes a page into semantic layout components and ranks them against the question with a late-interaction visual retriever adapted from page-level to within-page retrieval. Retrieval confidence then sets a question-specific evidence budget in place of a fixed Top-k. Finally, a reconstruction layer serializes the selected evidence for the downstream image processor, giving compatibility across VLM families while the detector and generative VLM remain frozen. Component-level adaptation raises within-page R@1 from 37.78% for the base ColSmol page-level retriever to 92.56%. Across nine frozen backbones, DocComp reduces visual tokens by 62.34–63.26% for Qwen with ANLS changes of −1.9 to +0.8 points, and by 33.20% for InternVL with changes of −4.4 to +3.0 points, improving three of four models. At similar reductions, it exceeds VScan and VisionZip by 14.14 and 16.92 ANLS points, and, counting its own detection and retrieval, cuts page-amortized compute by 61.04% on the reliable subset. These results support pre-VLM semantic evidence selection as an effective alternative to processing the full page first and pruning afterward.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.