Neural Language Dynamics: Short-Horizon Stability and Dissipation of Perturbations
Abstract
We study neural language dynamics: treating the text rollout as a dynamic system, quantized by a pretrained embedding model one token at a time. We formulate this inference-time process as a discrete input–state–output system in which tokens drive an internal Markov state, read out by the embedding model. Our benchmark spans 15 model configurations, 7 languages, parallel sentence-scale FLORES+, and document-scale FineWeb/FineWeb2 trajectories. Local geometry separates sharply by readout: last-token trajectories retain large turning and approximately constant token-scale speed, whereas pooled trajectories slow with prefix length. A predictive gate accepts a reduced or delay coordinate on 267 of 280 shards. Accepted causal coordinates are moderately predictable and usually contract rapidly under a fitted output-space surrogate; pooled coordinates are far more predictable but evolve near the unit timescale, while Chinese causal trajectories form a slower, more metastable intermediate regime. Paired interventions on English FineWeb across three Qwen3-Base sizes expose the central stability distinction. Under a shared suffix, plausible and random one-token edits contract with half-lives of roughly 1.5–5.5 tokens. Once the branches generate freely, their deviation instead returns to the ordinary sampling-noise scale, so shared-input contraction does not imply synchronization under free generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.