Probability-Weighted Pooling for Unified Context Compression and Reranking
Abstract
Retrieval-augmented generation (RAG) systems commonly use separate models for context compression and document reranking, requiring redundant encoding of the same query-document pairs. Unifying these stages allows a shared encoder to rank documents and prune the contents in a single pass. However, unified ranking requires choices in how to pool token representations and supervise the resulting ranking score. In this work, we introduce probability-weighted pooling, which weights document-token representations by the relevance probabilities predicted by the compression head. We theoretically prove that, as the predicted relevance weights in probability-weighted pooling become more accurate, its ranking score converges to that obtained using ground-truth evidence weights. We conduct comprehensive experiments across six benchmarks, comparing three pooling mechanisms and two ranking targets. Although the average improvement is modest, the gains are more pronounced on PubMed, where queries often involve multiple relevant documents, consistent with the statistically significant advantage observed in our aggregate comparisons. Document-level supervision yields a larger improvement over span-or-penalty supervision. These results highlight the greater impact of ranking supervision relative to pooling choice.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.