When Does Query-Aware Quantization Help? From Redundancy In-Distribution to Gains in Multimodal Retrieval
Abstract
Embedding-based retrieval at scale depends on vector quantization to hold large databases in memory. A natural idea is to compress vectors so that similarity to queries, rather than the vectors themselves, is preserved. Methods built on this idea, however, often yield only modest gains over strong query-agnostic baselines on standard benchmarks, and the field has lacked a principled account of when query awareness should help. In this paper, we find that on standard benchmarks the query and database distributions are so closely aligned that a rotation optimized on the data alone already captures the query structure, leaving query-aware objectives little to exploit. The redundancy persists even locally, where genuine query structure does exist: we show it adds nothing beyond local data statistics. On multimodal text-to-image retrieval (Yandex Text2Image), where the two distributions differ by construction, the situation reverses: at moderate compression rates, our proposed spectral method delivers 7.1 percentage points more Recall@10 than the standard OPQ+IVFPQ pipeline (a 26% relative improvement) and 1.7 points more than the strongest baseline we test, at equal memory and search cost; every query-aware and locally adaptive variant we test also improves on OPQ+IVFPQ in this regime. We distill both findings into a diagnostic that, from a sample of calibration queries alone, determines in minutes whether a given workload will benefit from query-aware quantization before any method is built.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.