Off-Policy Learning of Capacity-Constrained Allocations
Abstract
We study off-policy learning of capacity-constrained allocations from individual logged rewards under a shared linear model and an unknown logger that may assign zero probability to available actions. Capacity constraints determine which unresolved reward directions matter, captured by a source-radius modulus. We characterize the exact infinite-data identification threshold for a fixed source law and propose a minimax-regret allocation rule with high-probability worst-batch guarantees without overlap or minimum-eigenvalue assumptions. Marginal and integral outputs can have fundamentally different sample complexity. A constructed family requires samples for marginal outputs but for integral outputs. Experiments recover the predicted finite-sample behavior, and a semi-synthetic study on real LLM query contexts shows lower empirical failure and regret under source ambiguity for MRA than regression and robust baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.