AREA: Adaptive Relevance Evidence Allocation For Efficient LLM Reranking
Abstract
Large language model (LLM) rerankers provide strong relevance modeling but incur substantial inference cost because full-context inference propagates every input representation through all model layers. Existing efficient reranking methods typically rely on predefined compression configurations or externally selected operating points, despite substantial variation in where relevance evidence occurs and how much evidence different query–document pairs require. We formulate efficient reranking as adaptive relevance evidence allocation and introduce AREA, a pointwise reranker that learns both where full-resolution intermediate representations should be preserved and how much evidence each instance requires for deeper computation. AREA derives block-level localization supervision from the sensitivity of the relevance objective to compression, and instance-level budget supervision from counterfactual ranking feedback, without explicit token- or budget-level annotations. Under a fixed 50% retention ratio before ranking-guided optimization, AREA achieves 28.81 nDCG@10 on BRIGHT, compared with 27.72 for LTC. With adaptive budgeting, AREA reaches 30.21 nDCG@10 at an average retention ratio of 0.49, comparable to the separately optimized full-context variant at 30.10, while providing a controlled model-forward speedup. Consistent results on industrial search further demonstrate the effectiveness of task-derived localization and instance-specific budget allocation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.