acceptodds
Under review as a conference paper at ICLR 2027

Hidden Confounding Can Make CATE Validation Look Better as Causal Error Grows

Abstract

Because counterfactual outcomes are unavailable in observational data, conditional average treatment-effect (CATE) estimators are commonly developed and selected using observable held-out losses. A loss can rank candidates well without revealing whether any candidate remains causally accurate as the deployment environment changes. We uncover a systematic directional failure under hidden confounding: as confounding strengthens, causal error (oracle CATE root-PEHE) rises while held-out R-loss and a doubly robust pseudo-outcome loss improve. A population construction explains the direction: under hidden confounding, both losses reward movement toward the observed treated–untreated outcome difference rather than the CATE. Guided by this mechanism, we audit 24 structurally varied environments. Eight estimators spanning classical meta-learners, forests, orthogonal methods, and neural architectures exhibit this validation reversal in more than 92% of paired severity comparisons. Observable validation can still select a near-oracle-best member even as the best attainable CATE accuracy of the entire candidate pool deteriorates. We therefore propose severity-path validation audits as a new evaluation framework for causal-estimator development: developers should test whether observable model-selection criteria remain aligned with oracle causal accuracy as identification is progressively stressed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.