Generalization and Memorization along the Learning Trajectory of Neural Language Models: A Geometric Account of Categorization
Abstract
In this study, we demonstrate that backpropagation optimization updates the whole continuous landscape of the representation space from the beginning of learning. We train a series of decoder-only Transformers on controlled synthetic grammars and analyze changes in the representation geometry and models’ behaviour over the learning trajectory. We find that empty regions of the representation space are also systematically updated during early-stage optimization, even though these regions are not occupied by observed data points. This process progressively forms global geometrical structures in the whole continuous representation space. As a consequence of this optimization property, we find that early-stage learning is dominated by generalization, while memorization becomes increasingly dominant only at later stages of learning. From the earliest stage of learning, models progressively form categories that capture the true underlying data-generating process, rather than only memorizing specific input sequences. These categories are not simply sets of tokens, but continuous regions in the representation space. Model behaviour is driven by such categorization knowledge until the transition toward a more memorization-based representation. When memorization begins to dominate learning and representation, optimization in the representation space becomes increasingly localized around discrete observed data points. At the same time, the global continuous geometrical structures in the representation space are progressively destroyed. Consequently, model behaviour becomes less driven by category-based generalization and increasingly driven by the memorization of specific input data points. This learning trajectory contrasts with memorization-based theories of neural network cognitive representations. Our findings suggest that NLMs are inherently categorization-driven models: they achieve generalization by forming categories from distributional statistics, and this categorization behaviour begins from the earliest stage of optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.