acceptodds
Under review as a conference paper at ICLR 2027

Overcoming the KV Catch: State-Space Analysis of Transformer Generation Dynamics

Abstract

Transformer-based large language models (LLMs) are feed-forward networks, yet they are deployed autoregressively, giving rise to rich dynamical phenomena such as repetition loops. Analyzing these phenomena with dynamical systems tools, such as fixed points and their stability, faces a basic obstacle: due to their growing memory, these systems lack a fixed-size state. Here, we take a first step towards state-space analysis of growing-memory transformers. We show that an attention head without positional embeddings is exactly described by a non-linear self-interacting Markov chain with a fixed-size state, and use its timescale separation to derive a self-consistency condition for fixed points. Building on this condition, we develop an optimization method that finds both stable and unstable fixed points. Applied to position-ablated heads of a pretrained 1-layer transformer, it reveals stable repetition attractors, asymptotically unstable yet persistent semantic clusters, and saddle states between neighboring attractors. We then introduce a multi-timescale mean-field approximation for general attention layers, which closely matches the full model, and use it to find an unstable repetition whose predicted escape dynamics we confirm with real-model rollouts. Our work opens the door to applying dynamical systems tools to understand LLM generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.