acceptodds
Under review as a conference paper at ICLR 2027

Train Small, Guide Large: Optimization Scouting for Efficient LLM Post-Training

Abstract

Reinforcement learning with verifiable rewards (RLVR) can substantially improve LLM reasoning, but its computational cost increases rapidly with model scale. This work investigates whether low-cost RL on a smaller source model can replace full RL training for a larger target model. We find that transferable optimization signals already emerge at relatively early stages of source-model RL and can be used to derive effective parameter updates without costly iterative optimization on the target model. Building on these findings, we propose Optimization Scouting (OptScout), a framework for reusing RL-induced optimization signals across model scales. OptScout extracts evolving policy shifts during source-model RL and translates them into stable, target-specific update directions through low-rank gradient probing and cross-split consistency filtering. It then uses forward-only reward evaluation to select the best update direction and magnitude, and stops scouting once further source-model RL no longer improves transfer performance. Finally, the selected update is applied to the target model only once. Experiments across different source-side RL algorithms and target-model settings show that OptScout achieves higher average performance than direct target-side RL on multiple reasoning benchmarks, while providing up to a \(6.32\times\) end-to-end speedup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.