Pretraining with Masked Backstories in a Toy World
Abstract
Context-enhanced learning (CEL) involves augmenting context in large language models (LLMs) with special masked context to accelerate learning. It has previously only been explored using LLMs with billions of parameters during finetuning because CEL requires in-context learning (ICL) abilities to work. Here, we leverage a toy world (symbolically labeled randomly interleaved vector time-series from linear deterministic dynamical systems) that admits LLM-style “next token” pretraining and has been shown to exhibit multiple emergences of different ICL/recall abilities in tiny transformer models with mere millions of parameters. In this toy world, we can see a late transition from ICL to in-weights learning that also corresponds to a degradation of ICL performance. We enhance pretraining with additional masked context that allows the model to make near perfect predictions on the original training examples. Masking this additional context disincentivizes the model from memorizing it, and the capability of perfect prediction on the training example disincentivizes the model from memorizing the remaining portion of the training example. Not only does this enhancement suppress in-weights learning of the specific training systems, but we also show that it improves the model's performance in the seemingly unrelated task of associative recall. The study of this enhancement led to a few counterintuitive findings. First, despite such a model during training only seeing losses (and hence gradients) for tokens that are perfectly predictable, it generalizes well at test time when predicting tokens that are not perfectly predictable, nearly matching the performance of the optimal solution for those cases. Second, we found that memorization of 0.025% of the training systems was enough to degrade ICL performance on test data. Finally, when the enhancement is applied only at a specific position in each training example, the model still memorizes the training data, but can only recall these memories at the position where the enhancement was applied and not where the data was exposed to be memorized.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.