Better in Hindsight: Why Oracle Fusion Gains Are Hard to Learn and Where Evidence Helps
Abstract
Retrieval systems often fuse a fast retriever's score with a stronger verifier's score, and the fusion weight that ranks the correct item first varies from query to query. In hindsight, choosing the best of 21 weights for each CIRR query adds 10.2–12.2 R\@1 points over a tuned constant weight, whereas a ridge policy that predicts the weight from score statistics gains 0.32 points. We explain this gap with two exact identities. For R\@1, the successful weights of a query form an interval that can be computed from its scores, so headroom is exactly the share of queries whose interval excludes the constant. Every policy's net gain is its rescued failures minus its harmed successes. Across 21 vision–language verifier configurations on CIRR and FashionIQ and five BEIR text datasets, learned policies gain through caution: they rarely harm solved queries and seldom rescue queries whose successful weights lie far from the constant. The bottleneck is knowing which candidate is the target. A label-derived hint lets the same procedure recover a median 89% of headroom, competitor-aware features stop helping under an equal tuning budget, and only decoders trained to identify the target beat the tuned baseline. Acting on this diagnosis, the winner-path duel asks a second verifier about only the candidates that some weight can rank first, 1.6–2.0 calls per query. Under protocols fixed in advance, it gains 2.17 and 1.16 R\@1 points on new CIRR and FashionIQ queries, and 3.99 and 1.85 points (36% and 22% of headroom) when a larger verifier scores only those candidates. We recommend reporting reachable headroom, rescue and harm, and comparing adaptive policies under matched tuning budgets. 中文
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.