Can In-Context Learning Survive Recursive Pretraining?
Abstract
Can an in-context learner generate data from which its successors can relearn the same abilities? We study recursive pretraining of attention models on regression tasks, with each generation trained from scratch and evaluated on the original distribution. We distinguish optimization failure caused by small labels from structural degradation that survives energy normalization. For linear–quadratic tasks under population-context gradient flow, arbitrarily accurate single-generation learning coexists with a nonvanishing recursive clean-risk gap at every fixed training time. We derive a cubic threshold for preserving quadratic structure under polynomial training schedules. Alternatively, a calibrated signed combination of two attention-temperature views, followed by normalization, exactly preserves the task law and every generation's clean risk at fixed training time in the same population setting, without real-data refresh or parameter inheritance. We provide fixed-horizon SGD approximations and demonstrate improved retention in dense-attention experiments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.