STELLAR: Stage-Separated Testbed for Error Localization and Attribution in Retrieval
Abstract
Enterprise RAG pipelines are typically two-stage – an embedding retriever followed by one or more rerankers – yet almost universally evaluated with one end-to-end score that cannot say whether a deficiency originates in retrieval or reranking. We introduce STELLAR, a static benchmark built from 25,178 ServiceNow knowledge-base articles and 8,870 LLM-generated queries with mined, verified hard negatives, which scores retrieval and reranking independently – whole-corpus and fixed-pool modes, plus end-to-end – so a deficient score attributes to a specific stage. Every query also carries diagnostic metadata (lexical overlap, distractor hardness, length, product area, content type, LLM-assigned intent; test-retest ), disaggregating failure by cause, not just count. Across 18 model configurations, task formulation predicts quality more reliably than scale: a 1B cross-encoder significantly beats a 32B listwise LLM by 7.5 MRR points (), a 0.6B pointwise reranker matches or exceeds that same 32B model used listwise (directionally consistent but not itself significant at ), two same-family scaling curves are non-monotonic past moderate size, and cross-encoders dominate the quality-latency Pareto frontier outright. A diagnostic ablation surfaces an actionable failure mode: every reranker loses 10-14 MRR points on long source documents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.