WeRank: Scaling Listwise Reranking to Hundreds of Long Documents
Abstract
Large language models (LLMs) make listwise reranking appealing by enabling multiple candidates to be compared jointly under a shared query context. However, existing LLM-based rerankers are primarily designed for short passages and modest candidate sets, leaving listwise reranking over hundreds of long documents largely underexplored. To study this practically important regime, we introduce LoDoc, a real-world long-document reranking benchmark derived from production search traffic. LoDoc contains 14,884 queries, with 206.9 candidate documents per query and 2,081.7 tokens per document on average. We further propose WeRank, a scalable listwise reranking framework built on structured long-document compression. WeRank performs fine-grained chunk-level condensation and connects localized representations using a chained condensed chunk context scheme, enabling cross-region context modeling through compact memory states. The resulting representations enable hundreds of long-document candidates to be jointly compared within a shared ranking context. On the human-verified LoDoc test set, WeRank achieves 0.486 nDCG@10 and remains effective as the candidate pool scales to 300 documents. Under a matched representation budget, our structured compression design improves nDCG@10 from 0.400 with document-level compression to 0.443, a 10.8% relative gain. Experiments further demonstrate favorable inference efficiency and validate the key design choices of WeRank. We will make the dataset and source code publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.