acceptodds
Under review as a conference paper at ICLR 2027

Memorisation, convergence, and generalisation in generative models for text

Abstract

How to define generalisation for a generative model? According to an influential recent proposal, an image diffusion model is said to generalise if two models trained independently on disjoint subsets of a dataset generate nearly identical images given the same latent noise. Extending this idea to autoregressive transformers is challenging: unlike diffusion models, transformers lack a shared latent variable that naturally couples generation, and discrete tokens lack an obvious notion of overlap. Here, we introduce the survival probability – the probability that two transformers trained on disjoint subsets of the training data generate the same sequence under maximal coupling – and use it to empirically demonstrate a transition from verbatim memorisation to generalisation in transformers as training set size increases. We further show that independently trained transformers develop a shared predictive subspace, defined via the leading principal components of next-token probabilities, when trained on data sets that are substantially smaller than those required for convergence. We then introduce a simple auto-regressive generative model, the spiked Markov chain, and develop an asymptotically sharp characterisation of convergence and subspace recovery for this model. Our framework provides new tools for assessing generalisation in auto-regressive generative models and motivate investigating how transformer inductive biases shape both the emergence of shared predictive structure and the agreement between generated sequences.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.