ReAlloc: Hierarchical Optimization Budget Reallocation for LLM post-training
Abstract
Post-training is central to improving the reasoning capabilities of large language models (LLMs). While existing post-training methods differ in optimization objectives, they all require aggregating these objectives across sampled training data to form the batch objective. Standard practices typically aggregate objectives through averaging, with recent research generally prioritizing sampled data at a single granularity level of training units and tailoring to specific algorithms. In this work, we revisit objective aggregation as a hierarchical budget allocation problem over three granularity levels, i.e., prompts, rollouts and tokens, and identify signals that can guide the allocation of optimization weight at each level. At the prompt level, we connect outcome variance to the gradient of expected accuracy and find that prompts with mixed outcomes exhibit greater learning potential. At the rollout level, we show that incorrect trajectories closer to the centroid of correct rollouts preserve more coherent reasoning and exhibit less redundant gradient structure. At the token level, we find that local entropy prominence identifies reasoning transitions more precisely than entropy magnitude alone. Based on these findings, we propose **ReAlloc**, a hierarchical optimization budget reallocation framework that leverages these level-specific signals into aggregation weights over three levels while leaving the underlying objective unchanged. This design allows ReAlloc to be applied to different post-training algorithms. Experiments with GRPO and OPD across three models and seven benchmarks demonstrate consistent improvements, with only minor extra computational overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.