acceptodds
Under review as a conference paper at ICLR 2027

DocComp: Query-Conditioned Semantic Component Selection for Efficient Document VQA

Abstract

Document visual question answering (DocVQA) requires high-resolution inputs to read small text and structured content, but full-page processing turns even question-irrelevant regions into visual tokens. Existing visual-token pruning reduces the sequence only after visual processing, whereas aggressive region selection can lose answer-supporting context, and both can be tied to a specific VLM. To address this, we propose DocComp, a plug-and-play framework for query-conditioned semantic component selection that keeps only question-relevant layout regions before downstream VLM processing. DocComp decomposes a page into semantic layout components and ranks them against the question with a late-interaction visual retriever adapted from page-level to within-page retrieval. Retrieval confidence then sets a question-specific evidence budget in place of a fixed Top-k. Finally, a reconstruction layer serializes the selected evidence for the downstream image processor, giving compatibility across VLM families while the detector and generative VLM remain frozen. Component-level adaptation raises within-page R@1 from 37.78% for the base ColSmol page-level retriever to 92.56%. Across nine frozen backbones, DocComp reduces visual tokens by 62.34–63.26% for Qwen with ANLS changes of −1.9 to +0.8 points, and by 33.20% for InternVL with changes of −4.4 to +3.0 points, improving three of four models. At similar reductions, it exceeds VScan and VisionZip by 14.14 and 16.92 ANLS points, and, counting its own detection and retrieval, cuts page-amortized compute by 61.04% on the reliable subset. These results support pre-VLM semantic evidence selection as an effective alternative to processing the full page first and pruning afterward.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.