acceptodds
Under review as a conference paper at ICLR 2027

Proximal Policy Learning with Nonunique Outcome Confounding Bridges

Abstract

Learning treatment policies from observational data with unmeasured confounding requires valid comparisons of candidate policies. Proximal causal inference uses proxy variables and outcome confounding bridges for this purpose, but the observed bridge equation can admit multiple solutions. Existing results on nonunique bridges show that some causal functionals remain identified, yet they do not resolve policy choice when admissible bridges imply different rankings. In this paper, we characterize when policy value contrasts are invariant over the observed bridge solution set and, when they are not, formulate minimax regret over the joint bridge solution set. We derive an observed-moment dual and a convex empirical program for randomized mixtures of a fixed finite collection of policies. We establish a finite-sample excess-regret bound and give conditions under which a finite dual radius exactly represents the population objective. Under finite-feature rank and interior conditions, this exact representation yields root- excess worst-case regret at zero bridge residual tolerance for a fixed finite policy hull when the global optimization error is root-. We also relate worst-case regret over observed bridge solutions to worst-case regret over compatible causal models. Controlled experiments show the decision effects of selecting one bridge and relaxing joint restrictions, and quantify the tradeoff between minimizing worst-case regret and maximizing worst-case policy value.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.