CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
Abstract
Automating theoretical research requires generating candidate results and evaluating them reliably. Models keep getting better at the first, while the second remains hard. A common approach asks one large language model (LLM) to review what another produced, yet such reviewers are empirically unreliable: they may accept fabricated papers and catch the fabrication at close to chance rates (Jiang et al., 2025). We present CAUSALSMITH, a framework for automated theoretical research in causal inference built on the Lean proof assistant, where a proof is checked by a program rather than read by a referee. CAUSALSMITH rests on CAUSALEAN, a foundational Lean library for causal inference holding 10,575 machine-checked definitions and theorems, developed with language-model assistance under human design and review. Around it, we build a self-improving agentic pipeline that selects research topics, proposes results, formalizes statements, constructs proofs, and presents the resulting artifacts for human inspection. Moreover, the pipeline pairs Lean verification with a statement audit that compares each formal theorem against the informal claim behind it. We evaluate the system using artifacts produced by completed autonomous research runs. The source code, formal library, and run records are available at the anonymous supplementary replication package.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.