Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamcis, Convergence Rates, and Benefits of Off-Policyness
Abstract
We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted rollout data, with particular focus on a generalized version of classic REINFORCE — referred to as RE(S) — that updates the rollout distribution once every gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for REINFORCE and policy gradient methods, and on-policy sampling (i.e., a small value of , ideally ) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective even when the rollout distribution is updated infrequently, but theoretical understanding remains limited for the learning dynamics and convergence properties of these off-policy methods. To bridge these gaps, we develop a unified theory for RE(S) that covers the full spectrum of : it can be interpreted as a stage-wise iterative optimization process, where each stage takes gradient steps for minimizing the Kullback–Leibler distance to a fixed reward-weighted rollout distribution. For multi-arm bandits with softmax policies, our in-depth analysis and numerical experiments reveal three key findings: (1) for any fixed , RE(S) enjoys global convergence to the optimal policy as the number of rollout distribution updates , where denotes the total number of gradient steps; (2) we prove tight two-sided bounds for the convergence rate of RE(S), and show that its suboptimality gap achieves an asymptotic convergence rate, while only affects the length of a burn-in phase; (3) when initialized at a weak policy with a small optimal-action probability, RE(1) gets trapped around suboptimal policies for a long period, whereas RE(S) with a suitable avoids the detour and achieves significantly faster convergence to the global optimum, highlighting the benefits of off-policyness in this case.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.