acceptodds
Under review as a conference paper at ICLR 2027

Off-Policy Learning of Capacity-Constrained Allocations

Abstract

We study off-policy learning of capacity-constrained allocations from individual logged rewards under a shared linear model and an unknown logger that may assign zero probability to available actions. Capacity constraints determine which unresolved reward directions matter, captured by a source-radius modulus. We characterize the exact infinite-data identification threshold for a fixed source law and propose a minimax-regret allocation rule with high-probability worst-batch guarantees without overlap or minimum-eigenvalue assumptions. Marginal and integral outputs can have fundamentally different sample complexity. A constructed family requires samples for marginal outputs but for integral outputs. Experiments recover the predicted finite-sample behavior, and a semi-synthetic study on real LLM query contexts shows lower empirical failure and regret under source ambiguity for MRA than regression and robust baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.