When Evaluation Changes the Conclusion: Measuring the Reliability of Small RL Gains in Reasoning LLMs
Abstract
Reinforcement learning (RL) post-training has become a widely adopted approach for improving the reasoning capabilities of large language models (LLMs). However, many reported improvements are relatively small, raising an important question: whether current evaluation protocols provide sufficient resolution to reliably measure such incremental gains. We conduct a systematic empirical study of RL gain measurement reliability under different evaluation protocols. We evaluate Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct with RL-trained variants including Static-KL GRPO, GRPO, and No-KL using greedy decoding, stochastic sampling, multi-seed evaluation, paired bootstrap confidence intervals, and evaluation-size scaling analysis. Our results show that reproducibility across training seeds does not necessarily imply robustness across evaluation protocols. Under zero-shot greedy GSM8K evaluation, Static-KL achieves consistent improvements across three training seeds, whereas the same checkpoints yield near-zero gains under 4-shot evaluation. Under stochastic sampling with 16 generations, the measured RL–Base differences are only 0.2–0.5 percentage points, with confidence intervals including zero. Evaluation-size scaling further shows that larger evaluations provide more stable estimates of small effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.