Fixed-Budget Accounting Can Mislead Reinforcement Learning for Reasoning
Abstract
Reinforcement learning with verifiable rewards (RLVR) has made reasoning performance increasingly depend on how compute is allocated. At a fixed budget, should training spend compute on more optimizer updates, more parallel rollouts, or deeper refinement passes? This question is easy to state but hard to answer because the planned token budget may not match the number of tokens actually generated. We study this problem for Low-Rank Adaptation (LoRA) with Group Relative Policy Optimization (GRPO) under a strict fixed-budget audit across several LLMs. Our audit compares the three allocation axes under scheduled-token accounting, instruments realized-token usage, and then tests any accounting-driven reversal with direct realized-budget reruns. This exposes an accounting inversion: refinement can look stronger after realized-token normalization because its rollouts often terminate early, while the controlled scheduled-budget comparison favors spending compute on optimizer updates. The direct reruns resolve the ambiguity: the optimizer-update allocation still achieves performance across the evaluated benchmarks while using fewer generated tokens. We also run a matched inference-time audit comparing refinement depth with parallel sampling on reasoning benchmarks. Depth helps in some settings, but loses or ties in others, and out-of-distribution (OOD) evaluations show no consistent depth advantage. Our results turn a potentially misleading fixed-budget comparison into a reporting rule: state the budget being fixed, treat realized-token normalizations as diagnostics, and verify apparent reversals with direct reruns before making training claims.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.