RetVDQ: Retrieval-Aware Variable-Depth Residual Quantization for Visual Document Retrieval
Abstract
Multi-vector visual document retrieval (VDR) represents each page with multiple embeddings and scores it through MaxSim, which matches each query token to its most similar page embedding. This enables fine-grained retrieval but incurs a large storage cost. Existing compression mainly prunes or merges embeddings, but since different queries rely on different embeddings, removing any of them may discard needed evidence. Quantization keeps every embedding, yet existing quantizers only approximate the original vectors without preserving retrieval rankings, and assign the same precision to all embeddings regardless of their importance. We propose RetVDQ, a retrieval-aware variable-depth residual quantization framework that keeps every embedding and adapts its precision under a fixed budget. RetVDQ learns residual codebooks progressively so that every code prefix preserves the full-precision ranking, and then determines the depth of each embedding by the marginal retrieval utility of additional codes. Both stages are guided by synthetic queries generated offline from each page. The resulting index works with the frozen retriever and its original MaxSim scoring, requiring no model fine-tuning. Experiments on ViDoRe V1 and V2 show that RetVDQ consistently outperforms quantization baselines, and retains 99.3% of full-precision quality on average with only 3.2% of the original storage, far less than token reduction methods need for comparable quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.