acceptodds
Under review as a conference paper at ICLR 2027

Climbing the Causal Ladder: Twenty Language Models Share a Causal Variable

Abstract

Do different language models compute with the same variables? Universality is a founding hypothesis of mechanistic interpretability, yet cross-model evidence (similar representations, transferable steering vectors) shows that models are alike, not that they compute the same causal variable. We give a test that can establish sameness. It grades evidence by Pearl's ladder of causation and requires invariance across the behaviours that read a variable and across models, through maps that never see it. Applied to grammatical number in 20 language models from five families (70M–8B), it finds a causal variable shared across all of them, with evidence at every rung. Counterfactually, interchanging a direction estimated from labels alone at the subject noun mediates a median 0.85 of the full residual state's effect on five held-out agreement behaviours (0.86 in natural sentences), without moving gender. Interventionally, graded edits move every behaviour proportionally, and ablating the direction removes a median 0.53 of the number signal, against 0.01 for the gender direction. Across models, affine maps fitted only on generic text carry it with median efficiency 0.97 within and 0.95 across families (0.99 in natural sentences), while shuffled-pair maps, an untrained source and random directions carry nothing. The test also delimits and discriminates: the shared variable is the noun's own number, not number cued by a determiner (these sheep), and gender, causal in every model, is shared less reliably. Rankings of directions by behavioural strength, the field's usual evidence, reverse under calibration, recovery level and fitting cohort in our preregistered experiments. Invariance, not effect size, identifies the variables models share.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.