acceptodds
Under review as a conference paper at ICLR 2027

Language Models Need Sleep: Learning to Self Modify and Consolidate Memories

Abstract

The past few decades have witnessed significant advances in the design of machine learning algorithms– from early studies of task-specific shallow models to more general deep Large Language Models (LLMs). Despite showing promising results on tasks requiring instant prediction or in-context learning, existing models lack the ability to continually learn and effectively transfer their temporal in-context knowledge into their long-term parameters. Inspired by the human learning process, we introduce a "Sleep" paradigm that enables models to continually learn, distill their fragile short-term memories into stable long-term knowledge through replay, and recursively improve themselves through a "Dreaming” process. More specifically, Sleep consists of two stages: (1) Memory Consolidation: an upward distillation process, called Knowledge Seeding, in which the memories of a smaller model are distilled into a larger network to provide greater capacity, while preserving the knowledge. As a proof of concept, we present a new Generalized Distillation process for Knowledge Seeding, combining on-policy distillation with Reinforcement Learning (RL)-based imitation learning); and (2) Dreaming: a self-improvement phase, in which the model uses RL to generate a curriculum of synthetic data for rehearsing new knowledge and refining existing capabilities without human supervision. We provide rigorous theoretical analysis showing that the Knowledge Seeding objective admits a closed-form truncated SVD solution (under a linear upper network) in a space jointly shaped by the student’s on-policy covariance and a reward metric tensor. We also show that random MoE expert injection during dreaming provably improves the conditioning of the gradient covariance. Our experiments on long-horizon, continual learning, knowledge incorporation, and few-shot generalization tasks support the importance of the Sleep stage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.