acceptodds
Under review as a conference paper at ICLR 2027

Learning Rate Transfer for Reinforcement Learning Post-Training

Abstract

Reinforcement learning (RL) post-training of language models can be unstable and sensitive to hyperparameter choices, requiring extensive tuning that is particularly costly for larger models. We investigate whether learning rates selected on a smaller model can transfer to wider and deeper models during RL post-training. We show the Group Relative Policy Optimization (GRPO) objective is compatible with the CompleteP feature-learning scaling prescription. Across five verifiable-reward tasks and models with up to approximately one billion parameters, the scaled parameterization supports more reliable learning rate transfer than the standard parameterization, with a wider shared range of high performing rates across widths. These results support tuning on smaller models as a practical approach to reducing the hyperparameter search cost of RL post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.