acceptodds
Under review as a conference paper at ICLR 2027

Decomposing Decomposition: When and How Do Models Learn Combinatorial Structure?

Abstract

There is now abundant evidence that neural networks learn to decompose solutions to large problems by learning smaller subroutines for subproblems, which can often be reused and recombined, sometimes even facilitating helpful forms of compositional generalization. Comparatively little is known, however, about what drives this decomposition, and where we should expect to find it. The present work confronts this question by studying the dynamics of learning various *combinators*, that is, higher-order functions that determine how simpler functions combine. We first study which combinators can be learned by Transformer-based language models across model sizes and data configurations. We then find that the model's generalization accuracy on novel combinations is strongly correlated with its successful realization of the task's intermediate causal variables. We explain how the model identifies and learns these intermediate variables by tracing the gradients of training examples that share the same intermediate value yet have different downstream functions. We find that these gradients drive the clustering of the intermediate variable representations, and such clustering leads to both the emergence of the causal variables identified by prior interpretability research and the model's compositional generalization capacity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.