Policy Learning under Stable Rank Dependence
Abstract
Policy learning aims to personalize treatment decisions to improve population welfare. Yet population welfare maximization does not account for individual-level benefit and harm. Such quantities depend on the joint distribution of potential outcomes, while observed data generally identify only their conditional marginals. Existing approaches either optimize over separate worst-case couplings across covariate values or impose parametric restrictions on the unidentified dependence between potential outcomes, leading to conservativeness or strong structural assumptions. In this paper, we study policy learning when rank dependence is stable across covariates, allowing the conditional marginals to vary freely while leaving the common copula otherwise unrestricted. This restriction forces the least-favorable copula to be chosen jointly over the covariate distribution, yielding a sharper maximin policy value than rectangular partial identification. We derive the sharp value, characterize exactly when copula invariance changes the optimal policy, and develop a finite-sample learner that directly optimizes the resulting objective without estimating the unidentified copula. We establish uniform policy-value and regret guarantees and quantify robustness to approximate violations of copula invariance. Experiments confirm the predicted policy-switch behavior and show that the learner remains effective under moderate violations of copula invariance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.