Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
Abstract
Hidden confounding remains a central challenge in estimating treatment effects from observational data, as unobserved confounders can lead to biased causal estimates. Although recent work has explored the use of large language models (LLMs) for causal inference, existing approaches still rely on the unconfoundedness assumption and therefore cannot address hidden confounding. In this paper, we make the first attempt to mitigate hidden confounding using LLMs. We propose ProCI (Progressive Confounder Imputation), a framework that leverages two key strengths of LLMs: their ability to interpret rich textual information in causal benchmarks, which is often ignored by prior methods, and their internalized world knowledge, which enables the generation of plausible latent confounders and supports counterfactual reasoning. ProCI alternates between imputing a plausible hidden confounder based on the semantics of observed variables, and validating the imputation by empirically assessing whether the new confounder can help restore conditional independence between treatment and outcome. To enhance robustness, ProCI adopts a distributional reasoning strategy rather than direct value imputation, preventing collapsed outputs. Extensive experiments demonstrate that ProCI uncovers meaningful hidden confounders and substantially improves treatment effect estimation across diverse datasets, base models, and LLM backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.