Learning from Late Rewards: Asynchronous Data Synthesis with Realized Learning Utility
Abstract
Synthetic data is increasingly central to improving large language models, yet its value is commonly estimated from signals available before learning takes place. We argue that the true value of synthetic data is instead learner-relative and outcome-defined: it should be measured by the downstream improvement that training on the data actually induces. We call this quantity Realized Learning Utility. Grounding synthesis in realized utility, however, fundamentally turns data synthesis into a late-reward optimization problem: synchronous optimization preserves faithful feedback but blocks synthesis, whereas naive asynchronous optimization exposes the entire synthesis update to stale feedback from an evolving learner. We introduce LateSyn, a Predict-Early, Correct-Late framework that estimates future realized utility at generation time, applies a reliability-adaptive early update, and later corrects the original synthesis trajectory using the residual revealed by the actual learner outcome. A learning-utility memory further warm-starts prediction from previously resolved outcomes. Theoretically, LateSyn preserves the realized-utility target while reducing the synthesis signal exposed to delay from the full advantage to only its unresolved residual. Experiments across diverse agentic environments show that LateSyn produces substantially healthier learner trajectories in regimes where delayed feedback becomes harmful.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.