Transformers are Limited Causal Learners
Abstract
Autoregressive Transformers trained by next-step prediction underlie foundation models across modalities and appear to learn the structure of the worlds that generate their data. A growing body of work reads this structure as causal, through attention maps or Transformer-based causal-discovery estimators, yet when prediction recovers causal structure, and where it fails, remains unclear. In this work, we examine this question for time series from lagged structural causal models, distinguishing a variable's predictive parents, on which its conditional distribution depends, from its structural parents, which enter its generating mechanism. We show that under four identifying conditions the two sets coincide, so that structural parents can be read from the gradient of the log-density that a Transformer learns by next-step prediction, without any graph objective. Beyond these conditions, where causes are hidden, effects are instantaneous, or regimes are pooled, other past variables become predictive parents, and reading them as direct causes yields causal illusions that no model capacity can remove. Experiments with decoder-only Transformers bear out both halves: within the identifying conditions, the readout recovers structural parents across mechanisms, noise, dimensions, and lag horizons, including parents that act only on the shape of the distribution, while attention reflects this structure only in shallow models. Outside them, recovery declines, most sharply under hidden regimes, and supplying the missing information through the regime identity or dedicated causal-discovery methods serves as a remedy. Taken together, our findings support viewing Transformers as causal learners within the identifying conditions and as limited causal learners beyond them, and they locate the limits of causal learning from prediction in the information that the observed history carries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.