acceptodds
Under review as a conference paper at ICLR 2027

Seed Variance at the Stability Boundary Confounds Cross-Scale Comparisons in RL Post-Training

Abstract

Learning rates tuned on small models are often reused at larger scales in reinforcement-learning post-training. We ask whether the learning-rate optimum and the onset of instability transfer in the same way. Using GRPO on Qwen3-Base models at 0.6B, 1.7B and 4B parameters, we sweep learning rate and replicate configurations near and away from the transition to degradation. Across these scales, the 0.6B optimum remains near-optimal at 4B: the loss in final accuracy is 0.008 for the matched seed and 0.013 when the transferred setting is averaged over three seeds. The stability boundary is much less reproducible. Two grid steps below it, seed-to-seed variation is comparable to evaluation noise; at , the standard deviation rises from 0.007–0.010 to 0.116–0.160, a 15–20-fold increase relative to the reference seed variation. This variance coincides with a qualitative failure mode: runs abruptly lose the ability to terminate outputs, but the onset occurs at different training steps for different seeds. Consequently, a one-seed sweep produces a non-monotone scale trend that disappears under replication. On our grid, the practical distinction is sharp: a single run is enough to select a near-optimal learning rate, whereas measuring the stability boundary requires substantially more replication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.