Where Should a Fixed Budget of Human Data Go? Allocating Across Time and Modes in Recursive Training
Abstract
Recursive training on model-generated text degrades language models, and human-written data arrests that degradation. We study how to spend a finite human-data budget: one fixed lifetime stock of human-origin tokens, allocated both across recursive generations and across monitored regions of the human distribution. We formalise this as a constrained allocation problem and decompose it into four budget-matched policies that isolate each decision, spending randomly, on timing alone, on selection alone, or on both jointly. We enforce matching on two axes at once, lifetime human-origin tokens and total optimizer tokens, realising 0.0381% and 0.00% spreads against a 0.2% bound, and we assemble each corpus by displacement so every arm trains on identical data volume. We show that spending any human budget improves held-out regret by 4.04% over a control trained on that same volume, that concentrating the budget by mode is worth a further 9.59%, and that a back-loaded schedule improves endpoint regret by 6.51% on an interval excluding zero while leaving integrated regret unchanged, so the integral hides a timing effect the endpoint reveals. We then reconstruct every arm's realised purchases from the frozen manifests and verify them to the token against every archived chain total. That audit establishes which comparisons the executed run licenses: it shows the joint policy never left its base schedule, so the joint-versus-selection comparison measures how far two identical policies drift apart, and the value of coupling timing with selection remains open. We report 25 certified GPT-2 chains and 50 further chains on Pythia-160M and SmolLM2-135M as a descriptive replication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.