CHORUS OR ECHO? AN ASSOCIATION-AWARE RANKING LABEL MODEL FOR UNSUPERVISED RETRIEVAL FUSION
Abstract
Aggregating multiple retrieval sources is a natural strategy for scientific literature retrieval, where relevance judgments are expensive to obtain at scale. When several such sources agree, however, it is unclear whether their agreement should count as stronger evidence: sources may jointly promote relevant candidates or repeat the same mistakes. Recent evidence of correlated errors among large language models shows that stronger sources do not eliminate this ambiguity. We quantify the consequences of this ambiguity in a controlled experiment: increasing agreement on irrelevant candidates reduces equal-weight fusion's NDCG@10 from 0.39 to 0.17, despite unchanged individual-source ranking quality. We propose the Deliberative Noisy Label Model (DNLM), a label model for fusing retrieval sources without relevance judgments. DNLM combines Plackett–Luce-based learning of source reliability with pairwise associations estimated from score deviations relative to pair-excluded anchors to determine fusion weights. Across four scientific and biomedical benchmarks, DNLM achieves the highest NDCG@10 in four of five settings and the highest MRR in all five against fixed-rule fusion, individual sources, and weak-supervision baselines. Controlled analyses show that DNLM mitigates the loss from shared irrelevant agreement, and that disagreement matters as well: discarding the associations that record where sources deviate in opposite directions lowers NDCG@10 by up to 3.4 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.