Generalization Bounds and Representation Learning for Distribution-Valued Potential Outcomes
Abstract
Heterogeneous treatment effect estimation is central to understanding how causal effects vary across units with different observed characteristics. In many applications, however, each unit is observed through repeated post-treatment measurements that are more naturally represented by a probability distribution. Treatment can then change dispersion, tails, or shape even when the conditional mean is unchanged. In this paper, we study treatment-specific conditional Wasserstein barycenters for distribution-valued potential outcomes and define their conditional contrast through the difference of barycenter quantile functions. We develop a shared-representation estimator for learning these treatment-specific barycenters from observational data. We characterize the approximation induced by the learned representation, establish counterfactual generalization and finite-sample guarantees for target-population barycenter and contrast risks, and clarify how finite repeated measurements affect identification and estimation of the latent distribution-valued targets. Experimental results show that (i) shared fitting achieves lower mean contrast risk than separate fitting in all reported settings and lower treatment-specific barycenter risk in each synthetic family, and (ii) controlled experiments support the theoretically predicted effects of repeated measurements, quantile discretization, and representation learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.