Safe Policy Selection for Resource-Constrained Treatment Allocation with Batch-Level Underperformance Control
Abstract
Resource-constrained treatment allocation assigns a limited intervention jointly across each deployment batch. Once a baseline policy is in use, updated data and models often yield several candidate policies, making deployment a policy-selection problem. Mean gain alone is insufficient when a candidate can improve average outcomes yet incur a material loss on some deployment batches. In this paper, we study selection from a fixed candidate set under a constraint on the probability of batch-level underperformance relative to the baseline. This event is counterfactual and must be inferred from finite randomized calibration logs. We construct conditionally valid one-sided lower bounds for paired batch gains, convert them into conservative loss flags, and calibrate their frequencies with simultaneous exact binomial upper bounds. Selection then maximizes estimated mean gain among the certified candidates and the baseline. We prove finite-sample control of the selected policy's underperformance risk and characterize certification probability and the mean gain retained by finite calibration. Experiments with learned policies and controlled repeated-log and repeated-calibration designs illustrate the two sources of uncertainty and the resulting finite-sample behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.