acceptodds
Under review as a conference paper at ICLR 2027

RetVDQ: Retrieval-Aware Variable-Depth Residual Quantization for Visual Document Retrieval

Abstract

Multi-vector visual document retrieval (VDR) represents each page with multiple embeddings and scores it through MaxSim, which matches each query token to its most similar page embedding. This enables fine-grained retrieval but incurs a large storage cost. Existing compression mainly prunes or merges embeddings, but since different queries rely on different embeddings, removing any of them may discard needed evidence. Quantization keeps every embedding, yet existing quantizers only approximate the original vectors without preserving retrieval rankings, and assign the same precision to all embeddings regardless of their importance. We propose RetVDQ, a retrieval-aware variable-depth residual quantization framework that keeps every embedding and adapts its precision under a fixed budget. RetVDQ learns residual codebooks progressively so that every code prefix preserves the full-precision ranking, and then determines the depth of each embedding by the marginal retrieval utility of additional codes. Both stages are guided by synthetic queries generated offline from each page. The resulting index works with the frozen retriever and its original MaxSim scoring, requiring no model fine-tuning. Experiments on ViDoRe V1 and V2 show that RetVDQ consistently outperforms quantization baselines, and retains 99.3% of full-precision quality on average with only 3.2% of the original storage, far less than token reduction methods need for comparable quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.