Budget Allocation for PPO Text-Game Agents: Long Runs, Candidate Selection, and Successive Halving
Abstract
As reinforcement-learning post-training becomes a routine stage in producing language-model agents, one allocation decision recurs in every training campaign: with a fixed budget for delivering one model, should training go into one long run, or into several shorter runs from the same checkpoint followed by validation selection? Training several candidates creates the chance to select a better model, but under a shared budget each candidate receives less training, so the policies should be compared on the models they deliver. We compare a long run with two- and four-candidate allocations for Qwen2.5-1.5B and SmolLM2-1.7B under one PPO text-game recipe, with every comparison paired by block. A calibrated reference budget charges training together with the operational costs of producing and selecting a model, so all policies are compared at a common accounted budget rather than at equal elapsed time. Selection is sealed before independent evaluation, and every candidate of the two- and four-way splits is evaluated, so selection loss can be measured separately from the delivered gap. At the reference budget, splitting into two or four candidates reduces delivered success on the independent pool in all but one block-level comparison, and the registered reserved-pool check confirms the four-candidate deficit in both instances. On the independent pool, even the empirical best candidate of each four-candidate set falls below its paired long run in every block, so a different pick from those sets would not close the observed gap. With four candidates, the mean deficit exceeds ten percentage points in both instances. One successive-halving allocation improves substantially on four-way fixed splitting in both instances and on both pools. Doubling the budget improves the two-candidate strategy's relative performance in both instances. Its final ordering against the long run nevertheless differs, and the positive contrast in the first instance is sensitive to single-block deletion. These results identify a setting where selecting among more trained candidates fails to offset dividing training, while showing that the comparison depends on the allocation policy, budget, and model instance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.