Depth Is Not Enough: Compositional Generalization In Looped Transformers
Abstract
Compositional generalization requires reusing familiar operations in new combinations. For computations whose dependencies form a Directed Acyclic Graph (DAG), this involves both locating each operation's operands (syntax inference) and computing with them (syntax execution). Looped Transformers have demonstrated impressive generalization abilities by repeatedly applying a shared block, but it is unclear to what tasks they can generalize and why. In this work we introduce a task ladder, from prefix scans to serialized arithmetic trees, that progressively increases the difficulty of syntax inference. Models are trained on small compositions and tested on larger ones. The reference computation of every task is fully known, but models must infer its syntax from the input and then execute it. This lets us supply the true operands to a trained model and test the cause of a compositional failure. On this ladder, looped models fit every task in the training range but extrapolate only when each operand lies at a fixed offset from its operation. Larger models, other positional encodings, chain-of-thought and a pretrained 3.5B looped model do not close this gap. We establish syntax inference as the cause of this generalization failure through test time interventions: restricting a frozen model's attention to the true operand positions substantially improves extrapolation. Adding iterations, by contrast, does not help: beyond a few iterations, accuracy stops improving. Computational depth is therefore not enough: looped models must also learn to locate operands in a way that holds beyond the training range. Our results highlight limitations of current looped models and improve our understanding of what enables compositional generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.