Conformal Confidence Suppression for Offline and Online Learning-to-Rank
Abstract
Ranking systems often decide which candidates remain eligible before assigning display positions. Suppressing a useful candidate at this stage can prevent its recovery by the downstream ranker. We propose conformal confidence suppression (CoCo), a model-agnostic gate that suppresses only when a calibrated upper bound on incremental utility is negative. We distinguish eligibility from actual display, construct global, groupwise, and locally scaled gates, and prove finite-sample false-suppression control for an exchangeable calibration target. We also develop gated policy-value estimators and an online predict–gate–rank–update procedure, with conditional calibration results and a suppression-aware regret decomposition. On a public learning-to-rank benchmark, CoCo-M improves utility over the same ranker while empirical target violation tracks the chosen risk level. A sequential full-information simulation shows gains for evolving rankers, and randomized uplift data provide complementary calibration evidence. The results characterize how calibration and gate granularity trade off utility and conservative suppression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.