Skip, Adapt, or Route? Certified Deployment of Test-Time Adaptive Rerankers under Limited Labels
Abstract
Test-time adaptation of a re-ranker's query representation from pseudo-relevance feedback improves some queries and damages others, so its average effect does not decide whether to deploy it. We treat this decision as a choice among frozen actions, made from a small labeled sample of the target traffic: keep the incumbent ranking, adapt every query, or route adaptation with a label-free gate. We prove three results. First, a selector that sees only label-free signals has regret bounded away from zero on conditions it cannot distinguish, so target labels are necessary. Second, an exact value ladder bounds what any gate on a signal z can add over the better simple action by min(E[mu(z)+], E[mu(z)-]), where mu(z) is the conditional effect of adaptation; four anchor regimes determine its sign. Third, certified policy selection (CPS) replaces the incumbent only when a confidence bound on the paired improvement clears a margin. It has a finite-sample safety guarantee, an explicit regret rate that includes the cost of fitting a gate, a closed-form label budget, and a sequential variant. We replay stored per-query outputs of 23 benchmark conditions, a production search log, and a new self-adaptation testbed. With 50 labeled queries, uncertified selection deploys a harmful action in 23–70% of calibrations and an uncorrected t-test in 6–28%; CPS with a normal bound does so in under 1% on average, and with a finite-sample bound never, at the price of deploying rarely. The label budget is predictable from the effect size. Score entropy sees at most about a fifth of the per-query routing headroom, because adaptation mainly pulls the scorer toward the retriever, and recommendations from a fixed configuration reverse once step sizes and the retriever's own ranking enter the action set.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.