acceptodds
Under review as a conference paper at ICLR 2027

Fixed-Budget Accounting Can Mislead Reinforcement Learning for Reasoning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has made reasoning performance increasingly depend on how compute is allocated. At a fixed budget, should training spend compute on more optimizer updates, more parallel rollouts, or deeper refinement passes? This question is easy to state but hard to answer because the planned token budget may not match the number of tokens actually generated. We study this problem for Low-Rank Adaptation (LoRA) with Group Relative Policy Optimization (GRPO) under a strict fixed-budget audit across several LLMs. Our audit compares the three allocation axes under scheduled-token accounting, instruments realized-token usage, and then tests any accounting-driven reversal with direct realized-budget reruns. This exposes an accounting inversion: refinement can look stronger after realized-token normalization because its rollouts often terminate early, while the controlled scheduled-budget comparison favors spending compute on optimizer updates. The direct reruns resolve the ambiguity: the optimizer-update allocation still achieves performance across the evaluated benchmarks while using fewer generated tokens. We also run a matched inference-time audit comparing refinement depth with parallel sampling on reasoning benchmarks. Depth helps in some settings, but loses or ties in others, and out-of-distribution (OOD) evaluations show no consistent depth advantage. Our results turn a potentially misleading fixed-budget comparison into a reporting rule: state the budget being fixed, treat realized-token normalizations as diagnostics, and verify apparent reversals with direct reruns before making training claims.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.