acceptodds
Under review as a conference paper at ICLR 2027

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

Abstract

Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self-generated rollouts. However, we identify a critical support-limited bottleneck in this paradigm: on challenging reasoning tasks, the target model’s samples often exhibit semantic redundancy and pro- vide negligible reward contrast for policy updates. In this paper, we propose to overcome this limitation through a weak-to-strong learning paradigm, where a policy’s exploration is informed by a weaker but computationally efficient aux- iliary model. We introduce W2SPO, an off-policy RL method that injects short auxiliary segments—often as brief as 8 tokens—into intermediate target-model trajectories and the target model then completes the reasoning path from these diverted states. Policy updates are restricted to these short inserted segments based on final verifiable rewards. Empirically, W2SPO achieves superior perfor- mance among evaluated 4B-scale models on mathematical reasoning benchmarks, outperforming evaluated post-trained baselines. Compared with vanilla GRPO under the same sampling budget, W2SPO improves Pass@1 from 62.3% to 64.2% while achieving a 3.55× training speedup. Additional experiments show gains with a non-Qwen 7B target and on a multi-turn tool-use benchmark. These re- sults suggest that weak auxiliary branches can induce stronger target reasoning policies by expanding local exploration support. Code and data are available at https://anonymous.4open.science/r/W2SPO-1B86/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.