Fixed Rankings, Moving Leaders: A Relevance-Label and Source-Identity Audit of Scientific Retrieval Evaluation
Abstract
We show that a scientific retrieval benchmark can change its conclusions while every retrieval system stays fixed. In CER-Bench, a historical collection of 304 synthetic biomedical retrieval tasks over 4,936 PubMed records, replacing seed relevance labels with pooled automated judgments reverses nine of 45 pairwise system orderings on identical saved rankings (Kendall's τ_b = 0.584) and moves the mean-recall leader from a three-round search agent to a dense retriever. Label incompleteness is a known problem in retrieval evaluation, but it is rarely audited together with source identity: whether an identifier in the corpus refers to the article text attached to it. We audit both. An authoritative re-fetch of all 4,936 records traces a parser error that let cited-paper identifiers replace an article's own identifiers, changing 247 PMC identifiers, 226 of which occur in reference lists. A controlled Monte Carlo experiment with 1,000 paired, nested disclosures of the existing judgment pool locates the change of leader between six and nine revealed judgments per query, with paired empirical intervals overlapping zero at both points. We release a source-identity-checked retrieval dataset (4,936 documents, 10,313 chunks, 525 documents with conservatively restored own-article text), a hash-bound BM25 substrate, all saved rankings, judgments, disclosure arrays, and audit code. We deliberately do not release a corrected leaderboard: without human relevance validation or new model runs, the defensible output is an explicit boundary between reproducible arithmetic, verified source identity, and still-unvalidated relevance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.