Reinforcement Learning, On-Policy Distillation, or Both? A Finite-Budget Theory
Abstract
Reinforcement learning (RL) learns from rewards, while on-policy distillation (OPD) trains a student to imitate a teacher at states the student visits. RL updates can be noisy; OPD can inherit teacher errors. We ask when combining them improves on either method alone under a finite budget. We compare expected final regularized reward regret in a known nonlinear two-stage model, charging for rollouts, rewards, and teacher queries. Under explicit local conditions and bounded learning rates, a fixed mixture can outperform both pure methods even after independent rate tuning and additional updates funded by saved query costs. For the specified estimators, this guarantee covers all permitted RL schedules chosen before training and adaptive OPD rates and stopping. An explicit sufficient window relates the advantage to teacher error, reward noise, learner-dependent visitation, and budget. Retaining sampling variation is essential: average update directions can predict the wrong winner. A separate two-update result quantifies the benefit of adapting RL rates to observed progress and gives sufficient conditions favoring mixing or adaptive RL. Controlled simulations check these criteria. A controlled Qwen study provides complementary evidence: after independent tuning under equal query allowances, mixing has lower mean regret with noisy rewards, while RL has lower mean regret with clean rewards. Together, the results identify a benefit that cannot be explained by rate tuning or cost reallocation alone. Teacher guidance can help when rewards are noisy and teacher mismatch is moderate; reliable rewards, cheaper pure updates, or effective adaptation can remove that benefit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.