The Cost of Safe Skipping: Adaptive Routing with Paid Feedback
Abstract
Retrieval across providers with different search costs must decide both where to search and when to stop. Public routing and feedback-based stopping already address parts of this problem; we ask what reordering from paid returns adds once both are available. The key distinction is between finding the complete top-k and deciding to stop, since routing can affect these costs differently. We study this interaction with a neural maximum predictor, residual updates from paid returns, cost-sensitive acquisition and a learned stopping head. On 8.84 million MS MARCO passages, a frozen rank-eight policy reduces mean work by 6.59% relative to fixed order with the same predictor and stopping head. Its independent finite-pool omission-risk upper bound is 3.84% at a 5% target. Replay of 8,192 test queries locates the saving after result completion. Matched acquisition and threshold controls connect the gain to cost prioritization and stopping calibration, with one-shot cost ordering accounting for much of the reduction. On source-pure providers, earlier completion is offset by later searches. Warm neural execution retains the semantic throughput gain, while parallel static search has lower single-query latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.