From Static Compression to Query-Adaptive Representation for Efficient Visual Document Retrieval
Abstract
Visual document retrieval aims to efficiently identify fine-grained query-relevant evidence from long and visually rich documents. Multi-vector retrieval models are effective at localizing such evidence through fine-grained late interaction, but storing and matching hundreds of vectors per page incurs substantial memory and computation costs. Existing compression methods reduce this overhead through pruning or merging, yet the resulting page representations are typically fixed during offline indexing and reused for all future queries. This creates a fundamental mismatch between query-agnostic compression and query-dependent retrieval, since different queries may rely on different semantic aspects of the same page. We propose Query-Conditioned Reconstruction (QCR), a training-free framework that preserves compact reconstructable variation during indexing and instantiates query-specific page representations at retrieval time. QCR augments static compressed representatives with low-rank residual memory, selectively activates groups with high query-relevant reconstruction potential, and reconstructs query-conditioned vectors without accessing the original full multi-vector representation. A budgeted MaxSim-coverage objective further selects a compact set of complementary vectors for late interaction. Experiments across multiple retrieval encoders and benchmarks show that QCR better preserves Full-Index retrieval quality under aggressive compression, while improving evidence retrieval and downstream long-document question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.