acceptodds
Under review as a conference paper at ICLR 2027

Scheduled Information Value for Delayed Selective Labels

Abstract

Selective-label policy learning is challenging because current decisions affect both immediate payoffs and the outcomes available for future learning. This problem becomes harder under delayed feedback, where previously selected but unresolved outcomes may already provide substantial information in the near future. Because queue size alone does not reveal when such feedback will arrive, we then develop Scheduled-aware Information Value (SAIV) for sequential decision making under delayed selective feedback. SAIV combines a survival-based uncertainty baseline, which accounts for information expected from the pending queue, with a scheduled-information term that uses elapsed waiting times to capture timing information beyond a count-aware reference. This allows the policy to adapt exploration to both how much feedback is pending and when it is likely to arrive. Under a Beta-Bernoulli model, we derive the associated Bayes-risk reduction, establish a memoryless-delay null result, and bound sensitivity to delay-kernel estimation. Simulations and a sequential lending replay show consistent gains over exploration and delayed-feedback baselines, with mechanism experiments confirming that the improvements arise from exploiting the timing structure of unresolved feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.