acceptodds
Under review as a conference paper at ICLR 2027

Deep latents, long memory: Latent persistence explains temporal correlations in transformer activations

Abstract

The correlations between elements of a time series stem from the structure of the process that generated it. Here, we apply this lens to the internal representations of Large Language Models (LLMs): We define the autocorrelation function of the neural activity as the average dot product between the activations of a fixed layer at two context positions separated by a time lag. First, we find that the autocorrelations of deeper layers develop a slowly-decaying tail over training. We conjecture that this decay reflects the abstraction of the encoded features: Deeper layers encode more abstract variables, an abstract variable governs a longer stretch of the input, and its code persists over that stretch. We then test this hypothesis in two settings. In LLMs trained on synthetic data with known hierarchical latent structure, the persistence of latent codes predicts the qualitative structure of autocorrelations, including their development during training, and a large fraction of the empirical measures. In open LLMs trained on real text, we take sparse-autoencoder features as proxies for latent variables, and define their persistence time as the number of contiguous tokens where the feature is active. We find that the persistence time linearly correlates with the magnitude of autocorrelations across datasets and LLMs. In particular, the persistence time increases with depth as autocorrelations do, implying that deeper latent features persist for longer context spans. We complement this analysis by explicitly ablating the persistent sparse-autoencoder features and observing a consequent decrease in the autocorrelations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.