Sleep-Time Training: From Deployment Experience to Persistent Agent Improvement
Abstract
Large-scale pre-training and post-training produce capable LLM agents, but their parameters typically remain static after deployment. Existing deployment-time improvement methods mainly rely on additional computation during user interaction or external memory, which either increases user-facing latency or does not directly improve the agent's policy. We introduce Sleep-Time Training (STT), a deployment-time learning paradigm in which agents improve during user-idle periods. During user-facing sessions, the agent collects interaction trajectories. During sleep time, an auxiliary model converts these trajectories into corrected demonstrations and identifies defective intermediate actions. We first learn from replayed demonstrations with supervised fine-tuning, and then perform group-relative reinforcement learning over original and iteratively repaired trajectories. To study persistent improvement across repeated interactions, we introduce STT-Bench, a four-day benchmark containing 20 multi-day base task families and 20 held-out generalization families, with hidden user preferences and different levels of preference disclosure. Experiments show that STT improves future task completion and preference adherence without increasing user-facing inference cost, while retaining gains across multiple days and transferring to related unseen tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.