Dynamic Analysis of Large Language Models from a Basis Representation Perspective
Abstract
We introduce a basis representation framework for analyzing how the representation geometry of large language models evolves during training by characterizing the representation geometry of internal architectures through Gram matrices. By integrating gradient flow analysis, we identify three stages of pretraining: a focus stage in which representation geometry condense onto a low-dimensional structure, a dilution stage in which previously insignificant directions are activated, and a stabilization stage in which the representation geometry converges while the parameters continue to evolve. We empirically validate these predicted dynamics across multiple LLMs and datasets. In particular, we consistently observe the theoretically predicted dynamics when training open-source LLMs on real-world datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.