acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Causal Reasoning in LLMs: A Case Study in the Challenges of Procedural Generation

Abstract

Large language models (LLMs) appear to represent and reason about a large range of phenomena. In particular, recent studies have examined whether LLMs can accurately reason about causation. One of the most commonly used benchmarks for evaluating formal causal reasoning in LLMs is CLADDER, a framework for procedural generation of benchmark data sets that was released in 2023. Experiments with CLADDER appear to indicate that LLMs can perform formal causal reasoning in a large percentage of generated examples. In this paper, we examine what can be reliably inferred when LLMs perform well on test sets produced by CLADDER. We identify several examples of superficial and highly predictive patterns that the CLADDER generator embeds in data sets it generates. We find that even simple models such as random forests perform well on CLADDER-generated data sets after training on a sample of generated data. We also find that the performance of many commercial LLMs on CLADDER test sets is extremely brittle and that removing large sections of the text that are critical for formal causal reasoning has almost no effect on performance. Given these results, the most likely hypothesis for the strong performance of state-of-the-art LLMs on CLADDER is that they are exploiting superficial patterns learned from exposure to training data from the generator. Any inference that LLMs are exhibiting more complex causal reasoning must rule out this hypothesis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.