acceptodds
Under review as a conference paper at ICLR 2027

TRACE: Benchmarking Agentic Causal Estimation from Text under Covariate and Posterior Shift

Abstract

Average treatment effect (ATE) estimation is central to causal inference, but extending it to settings with open-ended text covariates is not straightforward. An estimator must connect information in individual narratives to a target-population effect, even when source and target populations differ in their covariates and causal mechanisms. Large language models (LLMs) can interpret text and perform mathematical reasoning, but those abilities alone do not ensure reliable effect estimates. To study this gap, we introduce TRACE-Bench, an end-to-end benchmark built from social narratives for evaluating ATE estimation across shifted source and target domains. Generic LLM agents, including direct estimation and planning-based approaches, are unreliable on this benchmark, while structured statistical estimators remain competitive. We therefore propose TRACE, a shift-adaptive framework that retains transferable source information and adapts the mechanisms that change. TRACE-AGENT learns and adapts natural-language contexts for causal estimation before doubly robust aggregation. TRACE-NN is an LLM-free, factor-augmented neural realization of the same framework. Across foundation models, TRACE-AGENT consistently improves ATE estimation over generic agent baselines, while TRACE-NN remains competitive with substantially larger LLMs. We finally prove that the proposed TARCE is doubly robust, which will give consistent estimation of ATE as long as either the propensity or outcome model is consistently estimated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.