Train Small, Guide Large: Optimization Scouting for Efficient LLM Post-Training
Abstract
Reinforcement learning with verifiable rewards (RLVR) can substantially improve LLM reasoning, but its computational cost increases rapidly with model scale. This work investigates whether low-cost RL on a smaller source model can replace full RL training for a larger target model. We find that transferable optimization signals already emerge at relatively early stages of source-model RL and can be used to derive effective parameter updates without costly iterative optimization on the target model. Building on these findings, we propose Optimization Scouting (OptScout), a framework for reusing RL-induced optimization signals across model scales. OptScout extracts evolving policy shifts during source-model RL and translates them into stable, target-specific update directions through low-rank gradient probing and cross-split consistency filtering. It then uses forward-only reward evaluation to select the best update direction and magnitude, and stops scouting once further source-model RL no longer improves transfer performance. Finally, the selected update is applied to the target model only once. Experiments across different source-side RL algorithms and target-model settings show that OptScout achieves higher average performance than direct target-side RL on multiple reasoning benchmarks, while providing up to a \(6.32\times\) end-to-end speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.