CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
Abstract
Large language models (LLMs) increasingly act as *data-science agents*, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis *or* data analysis benchmarks without causal data-generating structure or a range of causal data science tasks. Existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templatized variations rather than from systematic generation of novel synthetic causal structures. We introduce **CausalDS**, a benchmark for evaluating *causal reasoning in agentic data-science workflows*. Each benchmark instance is a *scene* consisting of a sampled structural causal model (SCM) with generated observational data and an accompanying synthetic natural-language story grounded in a real domain. We ground the *composition* of the benchmark components in empirical distributions obtained from real-world datasets, thus retaining empirical structure while reducing the memorization risk through completely synthetic generation. From each scene, we then derive tasks spanning all three of Pearl's rungs. Most tasks include a data science coding component, where the model typically needs to use several tools to arrive at the final answer due to imperfect observations. Abstention is tested directly: for targets the graph does not identify, the graded answer is to abstain. The benchmark thus jointly evaluates symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding. Because the generator is fully synthetic, the exam is not a fixed dataset: composition, difficulty, and size are parameters, and fresh exams can be drawn at will.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.