From Causal Diagnostics to Reporting Decisions: TrialVeto for OMOP Target-Trial Emulation
Abstract
Causal diagnostics expose mistimed treatment, inadequate overlap, negative-control falsification, and censoring instability in OMOP target-trial emulations, but convergence alone does not resolve whether an effect estimate supports an unqualified claim. We formulate reportability as a three-action decision—report, soft flag, or abstain—and introduce TrialVeto, a calibration-locked precedence rule with active-reason and estimand-drift traces. Evaluation scores every analysis in the declared workload by reporting coverage, risk among released estimates, risk among withheld estimates, and their difference. On 610 held-out drug–outcome pair–database analyses, TrialVeto emits 418 (68.5%) with median/P90 absolute log-HR disagreement of 0.12/0.31, compared with 0.19/0.58 for full-report Standard TTE and 0.15/0.43 after OHDSI calibration. The 192 abstentions have 0.34/0.71 error, yielding median/P90 gaps of 0.22 [0.16,0.29]/0.40 [0.29,0.53] and containing 75 of the 88 severe full-report disagreements. With thresholds re-locked, the median gap remains 0.24 on 153 unseen pairs and 0.25 under leave-one-database-out evaluation; at the identical 418-output budget, TrialVeto improves on the evaluated squared-error gradient-boosted selector by 0.025, 0.072, and 0.033 in median error, P90 error, and AURCC. Blinded clinical review recovers the intended reportability gradient (adjusted OR 10.4), and utility favors TrialVeto over OHDSI calibration whenever withholding costs less than 0.17 of a severe released error. In active-comparator, time-to-event OMOP emulations, TrialVeto turns post-estimation diagnostics into a transferable reporting decision with explicit error and availability costs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.