Retrieval-Aware Vocabulary Folding for Query-Side Compression in Asymmetric LLM Retrieval
Abstract
Asymmetric LLM retrievers reduce query latency by replacing online query-side LLM inference with a lightweight EmbeddingBag lookup. However, they retain a full tokenizer-aligned query vocabulary, introducing a large memory footprint despite the sparse usage of real retrieval queries. We reveal that this redundancy is not equivalent to token redundancy: rare tokens may carry critical lexical evidence for ranking, making naive vocabulary pruning unsafe. We propose Retrieval-Aware Vocabulary Folding (RAVF), a retrieval-aware vocabulary folding framework that compresses the query-side EmbeddingBag through utility-based anchor selection, token protection, and evidence-preserving folding. Instead of deleting removed tokens, RAVF maps them to compatible anchors while keeping the document encoder and retrieval index unchanged. Experiments on BEIR, CMTEB-R, and multiple LLM backbone scales show that RAVF significantly improves the accuracy–memory trade-off of asymmetric retrieval. Ablation studies further demonstrate that retrieval-aware folding and token protection are the main contributors, while lightweight InfoNCE calibration is sufficient without specialized distillation objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.