acceptodds
Under review as a conference paper at ICLR 2027

Be loud and relevant: pretraining of compositional tasks in linear RNNs

Abstract

Pretraining models with a curriculum of simpler tasks is a common approach to speed up training, but it is unclear what aspects of common task structure benefit learning, or how to choose a theoretically well-justified curriculum. Here we leverage recent learning theory in linear RNNs (Proca et al., 2025), and derive criteria for useful curricula that extend to compositional tasks. We find that pretraining tasks speed up learning and lead to more stereotyped solutions when they train model weights to large magnitude, while also building related structure to the target task. This mirrors results in feedforward networks, but does not rely on previously used statistical arguments of infinite-width networks. Instead we provide a mechanistic description for what constitutes a good CL task in RNNs: we demonstrate that factorized tasks representing the constituent parts of a composed task are ideal curricula because they possess hyperbolic learning dynamics in their pretraining phase that bypass saddles and slow points of the loss landscape.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.