acceptodds
Under review as a conference paper at ICLR 2027

Spend Now, Earn Later: A Bandit Framework for Delayed Resource Renewal

Abstract

Bandits with Knapsacks (BwK) captures sequential decision-making under hard resource constraints, but standard models either treat resources as non-replenishable or assume that replenishment is immediately usable. We study Delayed Replenishable BwK (BwDRK), a stochastic multi-resource bandit model in which each productive pull consumes resources immediately while the induced replenishment is credited only after a tagged bounded delay. Delayed renewal creates a physical funds-in-flight effect: resources can be sufficient in aggregate but temporarily unavailable for the next decision. We propose a queue-aware primal-dual algorithm with tagged queues, a bounded-delay primal wrapper, and arrival-order dual updates. The analysis first gives a pathwise liquidity identity that converts non-active waiting/terminal rounds into a maximum active-prefix deficit, and then decomposes expected LP regret into liquidity, primal-learning, and dual-control terms. Under , the statistical learning terms scale as up to resource-price constants. We also give a reserve-paced implementation based on the safe baseline, which enforces active-round slack and obtains complete sublinear expected regret against the corresponding reserve-aware comparator. The result separates the optimistic instantaneous-renewal LP from the operational price of maintaining lead-time liquidity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.