Linear Dynamical Modeling of Language Models
Abstract
How can we characterize the internal dynamics of a language model during generation? Modern LLMs produce text through high-dimensional, nonlinear hidden-state trajectories, making exact mechanistic understanding difficult. In this paper, we study these trajectories through linear dynamical systems (LDS), a classical model of state evolution whose eigenvalues provide interpretable measures of model memory and stability. Because recovering the high-dimensional full transition matrix or true latent state of an LLM is infeasible, we focus on a practical signature that approximates the dominant eigenmode of latent state evolution: the time-lagged covariance of observable states. We show that LDS-inspired summaries capture useful low-dimensional structure in both synthetic settings and real LLM generations. We further demonstrate that such summaries quantify generation stability across three downstream settings: evaluating the difficulty of safety-relevant jailbreaks, predicting mathematical reasoning correctness, and detecting hallucinations in retrieval-augmented generation. Most notably, classifiers built on these summaries predict answer correctness more accurately than frontier LLM-as-a-judge baselines (such as GPT-5.6 Terra), providing an interpretable and computationally efficient alternative requiring negligible additional inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.