Online Learning in Quota-Constrained Acquisition Games with Transient Impact
Abstract
We study strategic acquisition over a finite horizon when every agent must fill an exact quota, aggregate demand raises the price, and part of this impact persists before decaying. Our question is what an agent can learn from public prices when the competitors' quotas and actions are hidden. Under a natural stability condition, this game is strongly convex and has a unique pure Nash equilibrium. We show that for every fixed transient impact parameter , equilibrium schedules approach the uniform strategy and the price of anarchy approaches as grows. Across repeated episodes, the public price together with an agent's own schedule determine its cost gradient. Hence, decentralized online projected gradient descent achieves regret and mean-square last-iterate convergence without observing the number, quotas or schedules of the other agents. We obtain similar convergence guarantees for a Bayesian setting with (possibly correlated) private quotas. When the impact state is inherited across episodes, we characterize a self-consistent myopic equilibrium and its stochastic analogue. We also study learning within a single episode. We prove an external-regret lower bound against the best schedule in hindsight, which motivates comparison against finite families of bounded-memory scheduling rules. We present a model-free bandit algorithm and show that it achieves policy regret and empirical schedules approach a coarse correlated equilibrium at rate .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.