acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Reinforcement Learning: Scaling LLM Training with Self-Generated Data

Abstract

The scaling laws are a fundamental principle of large language models (LLMs), showing that model performance improves consistently as the amount of training data increases. In this paper, we propose Dynamic Reinforcement Learning (DynamicRL), a method that advances the scalability of RL for training LLMs autonomously. Unlike conventional approaches, DynamicRL does not require pre-collected datasets, human-labeled golden answers, or external verifiers for correctness; instead, it trains on self-generated data. DynamicRL operates by sampling data from the model as it evolves and using this self-generated data to optimize the model. Its dynamic characteristic allows the data distribution to continuously adapt to the evolving model, fostering better alignment between the training data and the model's capabilities. Experimental results demonstrate that DynamicRL can continuously improve model performance over more than a thousand training steps and achieve results comparable to models trained on large-scale external datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.