acceptodds
Under review as a conference paper at ICLR 2027

When Fusion Hurts: Diagnosing and Repairing Recoverable Failures in Heterogeneous Retrieval

Abstract

Heterogeneous retrieval systems combine channels such as lexical search, dense embeddings, visual retrieval, metadata search, and domain-specific indexes. Fusion is intended to preserve complementary evidence, but it can also suppress a relevant item that one specialist channel had already retrieved. We study this recoverable failure mode, which we call fusion hurt: at least one component ranks a relevant item within the target cutoff, but the fused ranking does not. Classical remedies such as tuned RRF, score calibration, quotas, and tree-based learning-to-rank repair some failures, but our analysis shows that they remain limited because they impose global or weakly query-adaptive aggregation policies over uncalibrated, source-specific evidence. We therefore use fusion hurt not only as an evaluation subset, but as supervision for a lightweight post-fusion listwise reranker. The reranker is trained over candidate-union rank, score, source, fusion, and disagreement features, with a query-level loss weight for fusion-hurt training cases; at inference time it uses no qrel-derived feature or fusion-hurt label. Across controlled and public heterogeneous retrieval settings, FH weighting improves the recoverable failure slice over the same plain reranker by +2.8 Recall@10 on VisRAG and +3.1 Recall@10 on MMDocIR, while BEIR text-hybrid validation shows a +3.1 Recall@10 and +5.7 nDCG@10 gain over raw hybrid overall. The experiments also expose where classical fusion remains insufficient and identify boundary cases where strong individual retrievers or tuned fusion already repair much of the failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.