acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Updates for Conformal Retrieval with Semi-Bandit Feedback

Abstract

Conformal retrieval constructs document sets at a prescribed coverage rate. Under semi-bandit feedback, a successful retrieval reveals the correct document, while a miss reveals only its absence. Existing conservative threshold methods provide guarantees for smooth coverage objectives, leaving the discrete number of returned documents unresolved. Coverage protection also changes the feedback used to learn the threshold. We propose counterfactual updates that recover whether the original proposed set would have contained the correct document. Whenever another miss would exceed the cumulative miss budget, the method returns the full candidate pool and uses the revealed document to evaluate the original proposal. Assuming every candidate pool contains the correct document, this preserves the unprotected learner's threshold trajectory and guarantees cumulative coverage after every query. For i.i.d. queries with fixed scores, a unique continuous target quantile and continuity of expected document cost at that quantile, average returned-set size converges to the best feasible fixed-threshold cost. In contrast, updating from protected hits keeps the threshold at or below its initial value and requires full-pool returns on a positive fraction of queries when initialized below the target quantile. The decreasing-step-size learner tracks that quantile directly, while conservative estimation certifies a threshold through a distributional bound. In the initial ten-order SQuAD comparison, our method returns 2.64 documents per query versus 9.44 for conservative threshold estimation. Protection adds about 0.09 documents per query relative to the same unprotected learner with identical step sizes, while maintaining at least 90% cumulative coverage after every query. With unchanged parameters, evaluation on Natural Questions also produces smaller sets than the compared conservative and constant-step methods. These results show that learning from the original proposal supports efficient retrieval with cumulative coverage protection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.