Validating Causal Benchmarking in Clinical Registries: Antimicrobial Timing in Septic Shock
Abstract
Generative causal benchmarks offer a way to compare treatment-effect estimators when the true effects in observational data are unknown. Their appeal is to combine known, model-defined effects with data that resemble the original study. However, a benchmark can pass verification simply because the tests cannot detect its errors. We investigate this problem through controlled experiments and the Cooperative Antimicrobial Therapy of Septic Shock (CATSS) registry of 11,187 patients from 32 hospitals, where we study antimicrobial timing and mortality. Septic shock is a time-critical condition with substantial mortality, and effective antimicrobial therapy is often initiated under urgent clinical uncertainty. This makes treatment timing a consequential setting in which reliable causal evaluation is particularly important. We find that outcome type sharply changes how added covariates affect verification power. Under severe generator error at , adding just five covariates reduces the energy test's detection rate from to for binary outcomes. In the corresponding continuous-outcome setting, detection remains through 50 covariates. A similar failure appears in CATSS: energy and maximum mean discrepancy detect a deliberately broken generator using treatment and outcome alone, but miss it once patient covariates are included. A classifier two-sample test detects the failure in both representations. On benchmarks generated by two model families, estimator comparisons further reveal robustness differences obscured by full-model evaluations: the doubly robust estimators remain comparatively stable when either the treatment or outcome model alone is restricted. These comparisons inform our clinical analysis, where effective antimicrobial therapy within one hour of shock onset is associated with a 14.4-percentage-point lower adjusted mortality risk. These findings show that apparent agreement between generated and observed data is not enough. Before using a benchmark to judge causal estimators, we should establish which generator failures its verification procedure can actually detect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.