acceptodds
Under review as a conference paper at ICLR 2027

Critical Batch Size for LLM Policy Optimization

Abstract

Policy optimization has become the dominant post-training technique for large language models, but it is computationally expensive. In pretraining and supervised learning more broadly, the critical batch size (CBS), which is the batch size beyond which further increases yield diminishing returns in training steps, is well studied and guides how compute is parallelized at scale. No analogous understanding exists for policy optimization, where the question is harder: the objective changes with the policy, sequences generated from the same prompt are not independent, and the batch size itself decomposes into the number of prompts and the number of rollouts per prompt. We develop a theoretical and empirical framework to study the CBS of policy optimization, focusing on GRPO in the on-policy regime. Theoretically, we present a noise model that decomposes the gradient noise of GRPO into a prompt-level and a rollout-level term, and predicts that adding prompts reduces both while adding rollouts reduces only one. Empirically, for GRPO on mathematical reasoning, we observe near-linear reductions in optimizer steps up to 16K sequences per step when scaling prompts. Together, we give concrete guidance on scaling prompts, rollouts, and learning rate, while shedding further light on GRPO batch scaling dynamics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.