Differential Neural Reranking: Updating Rankings Without Reranking Everything
Abstract
A revised query may need a new ranking, but not a rescore of every candidate. We ask how few candidates must be rescored to reproduce the top-k of a full rerank, and propose differential neural reranking, which uses the previous ranking to choose them. Across eight retrieval domains and two reranker architectures, revisions shift nearly every candidate's score. Much of this movement is shared across candidates and preserves their order, while the candidate-specific changes that alter the top-k concentrate near the cutoff. On 273 successive session revisions, an exact oracle needs only 2.65 of 64 fresh evaluations on average to recover the ordered top-10. A two-parameter model predicts cutoff crossings across rerankers and benchmarks without refitting, capturing the rank structure that determines which candidates require fresh evaluation. On fixed candidate pools, differential reranking speeds up the reranking stage by about 6× on GPU at fidelity nDCG@10 above 0.99. The effect persists in end-to-end pipelines, with savings large enough to offset the added scoring cost as the candidate pool changes. On 250 held-out TopiOCQA conversations, with the policy and evaluation protocol fixed in advance, differential reranking reduced mean end-to-end latency by 14.40%, outperformed a BM25 baseline with the same reranker budget, and kept fidelity nDCG@10 at or above 0.98 on every later turn in 249 of 250 conversations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.