From Memorization to Length Generalization: Grokking in Recurrent Models
Abstract
Recurrent models can in principle process arbitrarily long sequences by reusing the same state transition, but fitting short sequences does not guarantee that the learned transition will generalize to longer ones. We study this gap in an overparameterized linear RNN trained on sequential modular addition. We show that zero training loss alone is insufficient for length generalization because decoder-invisible hidden directions can remain unconstrained. In contrast, sufficient hidden-state coverage forces every zero-loss solution to implement the correct reusable recurrence. Combined with a small-initialization convergence result, this gives a tractable setting where exact arbitrary-length generalization can be characterized. We then explain why length generalization may emerge much later than short-horizon fitting. Near a generalizing solution, the delay is controlled by slowly decaying training modes and by their visibility at different rollout horizons. For decoder-invisible directions, we derive the local curvature and link the slowest learning rate to the conditioning of hidden-state coverage. This provides a quantitative mechanism for length grokking: recurrent errors can become negligible on short rollouts while remaining significant on longer ones. Our results connect hidden-state coverage, optimization dynamics, and the emergence of reusable recurrent computation from finite-horizon data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.