acceptodds
Under review as a conference paper at ICLR 2027

Sample Efficient GRPO via Learning to Reuse Tokens Across Trajectories

Abstract

We propose FORK-RL: forking rollout trajectories in the middle and reusing the corresponding prefixes to boost RL post-training efficiency of LLMs in coding. GRPO estimates advantage based on group variety, hence the learning signal vanishes when all solutions in a group agree, whether they are all correct or wrong. As a result, such generated groups waste rollout budget and limit performance. Rather than increasing the sheer group size, FORK-RL resumes from the middle of existing trajectories to generate new solutions with different outcomes without additional rollout budget. We further introduce a learnable light-weight score head to identify the critical forking tokens. The score head decides where to fork according to the hidden states of random candidates. We train this head with a fiip reward when the new soluton leads to a different outcome. Our method not only endows the disagreement within groups, but also achieves finer grained credit assignment in the prefix-sharing trajectories. On competitive programming benchmarks, we achieve better performance than the compute-matched baseline: on CodeContests with a 4B policy model, a 17.8% relative gain in pass@1 (24.5% against the baseline’s 20.8%). Notably, our method reaches the baseline’s training accuracy on 62% fewer decoded tokens at 4B and 46% fewer at 8B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.