acceptodds
Under review as a conference paper at ICLR 2027

Q-Loss: A Training Objective for Direct CATE Optimization Across Meta-Learner Families

Abstract

Conditional average treatment effect (CATE) estimation increasingly relies on model-agnostic meta-learners that decompose the problem into standard supervised-learning subproblems. These learners, however, are trained on surrogate losses only loosely aligned with treatment-effect quality, and a recent large-scale benchmark finds that 62% of fitted CATE estimators are degenerate—no better than predicting no effect at all. The Q-statistic of Yu & Sun (2025) provides an unbiased estimate of CATE risk that requires no counterfactuals and has been used to evaluate and rank estimators. We develop a differentiable, end-to-end algorithm for training meta-learners directly on this statistic—with learned feature gates that select treatment-effect modifiers—and instantiate it across the S-learner, X-learner, and DML-Learner families. On real and synthetic randomized-trial benchmarks, we then characterize when this helps: the apparent failure of the raw inverse-propensity target is largely a matter of propensity misspecification, which estimating the propensity and adopting a doubly-robust target resolves; and the resulting gains reflect genuine improvements in treatment-effect quality rather than the training statistic itself. Together, these analyses clarify when direct optimization of the Q-statistic is preferable to standard baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.