Group-Aware Verification Scheduling for Reinforcement Learning with Execution-based Rewards
Abstract
Reinforcement learning post-training of LLMs on tasks such as coding computes each reward by executing the rollout's output, for example by applying a generated patch in the task's environment and running its tests. Verification time can vary by an order of magnitude across tasks, and verification is failure-prone. Group-based algorithms such as GRPO make this worse, because a group of rollouts can enter training only when every member has a result, so the learner waits for the slowest member of its slowest group. Existing systems give no priority to the verification jobs that would complete a group. We present LastCall, a runtime that schedules verification for group-based reinforcement learning. It bounds each round of verification with a waiting window, after which incomplete groups keep verifying and can join a later batch whole under a one-update staleness bound, and gives each free verifier worker to the group closest to completion. Implemented inside asynchronous veRL, LastCall cuts model update time by 13% to 28% across three model scales in our experiments, without lowering the model quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.