acceptodds
Under review as a conference paper at ICLR 2027

CompRank: Efficient Long-List LLM Reranking via Token-Level Compression and Decoding-Free Scoring

Abstract

Large language model (LLM) rerankers have become an important component of modern retrieval and retrieval-augmented generation pipelines, yet their computational cost makes reranking long candidate lists expensive. We introduce CompRank, a token-efficient reranking framework that reduces redundant computation through contextualized KV compression and decoding-free attention scoring. CompRank first encodes each document with its full causal context and then exposes only a small subset of the resulting key–value (KV) states to the scoring stage, allowing the retained states to preserve information from the omitted context. We instantiate CompRank in two complementary variants: CompRank-QC, which conditions document representations on the query for stronger relevance modeling, and CompRank-Cache, which uses query-independent document representations to enable cross-query KV reuse. Experiments on seven BEIR datasets show that CompRank-QC achieves an average NDCG@10 of 46.7 with Mistral-7B, outperforming the evaluated efficient reranking baselines, with consistent effectiveness gains also observed on Qwen2.5-7B. Notably, with Mistral-7B, Step-10 compression retains only 10.2% of document KV states while reducing average NDCG@10 by only 0.5 points compared with full-token scoring. It also yields approximately end-to-end speedup at 500 candidates even when documents are re-encoded for each query, while document-side KV reuse further reduces online latency by amortizing document encoding across queries. On TREC-COVID, CompRank further generalizes from 30-document training lists to candidate sets of up to 500 documents. These results demonstrate that contextualized KV compression and reusable document representations provide an effective path toward scalable long-list LLM reranking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.