Learning Rate Transfer for Reinforcement Learning Post-Training
Abstract
Reinforcement learning (RL) post-training of language models can be unstable and sensitive to hyperparameter choices, requiring extensive tuning that is particularly costly for larger models. We investigate whether learning rates selected on a smaller model can transfer to wider and deeper models during RL post-training. We show the Group Relative Policy Optimization (GRPO) objective is compatible with the CompleteP feature-learning scaling prescription. Across five verifiable-reward tasks and models with up to approximately one billion parameters, the scaled parameterization supports more reliable learning rate transfer than the standard parameterization, with a wider shared range of high performing rates across widths. These results support tuning on smaller models as a practical approach to reducing the hyperparameter search cost of RL post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.