Large Language Models as Proxy Designers for Proximal Causal Inference
Abstract
Proximal causal inference (PCI) identifies causal effects under unmeasured confounding using two proxies for the latent confounder: a treatment control and an outcome control . Its practical bottleneck is the step before estimation, namely partitioning the covariates into , and observed confounders so that the identification conditions hold, a judgment that needs fluency in both PCI and the application domain. We show that large language models (LLMs) are increasingly capable of making this judgment. From the PCI conditions, a brief study context, and a description of each covariate, an LLM proposes what the latent confounder represents and assigns every covariate a role; a stratum-targeted outcome-bridge objective then yields conditional effects from one global bridge. We evaluate on three real cohorts with external checks. On a 10M-sample food-delivery corpus with a separate randomized price experiment, three LLM partitions give – error (relative to the randomized experiment), two of these are below and one is tied with the best baseline. In the Framingham Offspring study, every LLM partition that admits an estimate lands inside the range randomized trials report, where a third of random partitions and the published proximal split do not. In the SUPPORT study, all eight LLMs give estimates within the published confidence intervals and one reproduces the expert allocation exactly. For the (cardiovascular) Framingham and SUPPORT datasets, we also compare LLM covariate partitions to those from multiple Board-certified cardiologists.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.