acceptodds
Under review as a conference paper at ICLR 2027

Your Best Retriever May Not Be Best for Your Reranker: Adaptive Coordination for LLM Recommendation Cascades

Abstract

Large-scale recommender systems typically adopt a two-stage retrieve-then-rerank cascade, in which a retriever narrows a massive item pool to a small candidate set, and a reranker produces the final ranking. This creates an objective misalignment between recall-oriented retrieval and precision-oriented reranking: the retriever aims to ensure that relevant items are included, whereas the reranker must accurately distinguish them from competing candidates and place them at the top of the final ranking. Retrieval should therefore provide not only relevant items, but candidate sets that the downstream reranker can reliably resolve into high-quality final rankings. Prior work often addresses this misalignment through joint optimization, but discrete candidate selection limits end-to-end gradient flow, while the retrieval and reranking objectives may induce competing optimization directions. We propose R4Rec (Retrieve–Rerank–Reallocate–Refine), a cascaded framework that uses two trainable feedback mechanisms to coordinate pre-trained retrievers and an LLM-based reranker, aligning candidate generation with final reranking without updating the underlying retriever or reranker components. Across-Sample Reallocation learns context-dependent retriever quotas, assigning larger budgets to retrievers whose candidates better support the downstream reranker. It derives each retriever’s contribution from closed-form, reranker-aligned Shapley values and uses this feedback to train a contextual bandit. Within-Sample Refinement instead operates at the request level: a separate contextual bandit decides whether to stop or refine the current candidate set and, when refinement is needed, generates a diagnostic query to retrieve additional candidates before reranking. Experiments on Amazon and MovieLens show that R4Rec improves accuracy at practical inference cost, with a statistically significant 6.43% relative CTR lift over the production baseline in an online A/B test.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.