acceptodds
Under review as a conference paper at ICLR 2027

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

Abstract

We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates both whether an agent can solve a problem using causal evidence and whether its answer is grounded in a faithful recovered causal mechanism. Each episode hides a randomly sampled structural causal model (SCM) that no prior knowledge supplies: the agent reads prior records, intervenes on a manipulator crystal, and predicts the held-out frequency of a reactor crystal built from the same SCM. Because the reactor exposes every property except frequency, the prediction needs only that variable’s parents and coefficients, so we score this Target Mechanism Recovery separately from Full System Discovery. Experiments show that correct prediction does not imply Full System Discovery. Across matched controls on functional form, hidden perturbations, and graph structure, the two levels move independently or in opposite directions. From ten observations alone on 6-node graphs, GPT-5.2-high recovers the parents of frequency exactly and answers 86% of reactors correctly, yet reaches 0.471 all-edge F1. The same ten rounds spent on interventions lift all- edge F1 from 0.47 to 0.82 for ten accuracy points. Failures come from premature commitment, not exhausted budget, and a consistency check recovers part of the loss. A shift control on a quarter of the agents’ budget recovers every main- suite graph exactly, so the gap measures experiment design. CausaLab therefore separates the two levels and exposes current LLM agents’ limits as experimental causal reasoners. We release the full SCM dataset and benchmark code to advance future research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.