acceptodds
Under review as a conference paper at ICLR 2027

Understanding Retrieval and Reasoning in LLM Reranking under Availability Bias

Abstract

Large language models are increasingly used to rerank candidate items in recommender systems, where they may rely on both reasoning and retrieved evidence. We show that retrieval is not merely auxiliary to reasoning. Across two recommendation benchmarks, a small set of selected reviews matches a long reasoning trace on HR@1 at substantially lower inference cost, and the two signals combine. However, evaluating retrieval in this setting is confounded by a previously underappreciated bias. LLM rerankers systematically shift probability toward candidates that are accompanied by evidence-like text, even when that text belongs to another candidate or carries no useful information. We call this *availability bias*, by analogy to human judgment, without assuming a shared mechanism. Matching only the total number of retrieved passages does not remove the bias: when evidence is allocated unevenly, the allocation itself changes the ranking. We therefore introduce matched controls that separately account for evidence presence, count, allocation, content, and attribution. Under these controls, genuine evidence selection improves HR@1 by -, while much of the apparent advantage among common retrieval heuristics is explained by where they allocate evidence rather than by which passages they select. Our results show that retrieval is a first-class component of LLM reranking, but also a potential source of hidden ranking manipulation. Reliable evaluation must control for evidence availability and attribution, or retrieval systems may receive credit for effects unrelated to the information they retrieve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.