Composition and Closure: Causal Tests of Linguistic Operators in Transformers
Abstract
A causal account of grammar in a Transformer must explain not only whether individual linguistic operations can be realized, but whether those operations continue to act correctly on states produced by earlier operations. We formalize this requirement as complexity-bounded causal realization: rank-bounded interchange maps are learned from unary counterfactuals at a fixed layer and span, then evaluated under held-out composition and on their intervention closure, the states reach- able through valid intervention sequences. We show formally that perfect unary correctness on observational states does not imply correct composition, and give conditions under which closure-level one-step control bounds error along composed interventions. To study this distinction empirically, we introduce OpClosure, a synthetic benchmark with exact counterfactuals and controlled operator depth. Across independently trained Transformers, maps that perform well on unary interventions degrade systematically when applied to closure states, including after an earlier operation has been realized correctly. A modular reference is substantially more robust, while training with composed targets recovers much of the lost closure-level performance. We further observe analogous failures in controlled English experiments across pretrained models from multiple families and scales. These results suggest that unary causal success is insufficient evidence for compositional grammatical mechanisms: causal realization should be evaluated on the states that the proposed mechanism itself produces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.