EVERY BYTE MUST RANK: A RETRIEVAL-AWARE CODEC FOR FROZEN MULTIMODAL EMBEDDINGS
Abstract
Multimodal retrieval systems increasingly store every candidate as a high-dimensional embedding, and the bill is paid per vector: one billion 2,048-dimensional FP32 embeddings occupy 8.2 TB before any index is built. This scale makes embedding storage a critical systems problem. Yet the tools used to compress these vectors were designed to reconstruct them. Product quantizers minimize reconstruction error even though retrieval depends on the resulting rankings, and fixed-width codes charge every symbol equally regardless of predictability. We introduce EmbedCodec, a codec trained for ranking rather than reconstruction. A lightweight adapter shared by queries and candidates is optimized through the hard product-quantizer assignments used at deployment, reshaping the frozen encoder’s geometry for post-quantization retrieval. A task-conditioned Gaussian mixture models the resulting indices through codeword geometry; rANS then entropy-codes those indices, and rate is measured from serialized bytes rather than nominal bit width. Results on five retrieval tasks spanning text, cross-modal, and visual-document search with two frozen encoders show that EmbedCodec uses about 125 B/vector, a roughly 65× storage reduction relative to FP32, while retaining 99.6–99.8% of the corresponding five-task average NDCG@10. At about 500 B/vector, it exceeds the FP32 average on both encoders. Every persisted stream is exactly decodable back to its input index tuples, and each of the ten encoder–task panels contains at least one Pareto-nondominated EmbedCodec operating point. Our findings suggest that retrieval-aware source and entropy coding can substantially reduce embedding storage without requiring faithful coordinate reconstruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.