Heuristic-guided Branching-Replay Group Relative Policy Optimization
Abstract
Post-training has proven highly effective in improving the reasoning capabilities of large language models (LLMs). However, growing evidence suggests that reinforcement-learning-based post-training encounters a pronounced performance ceiling on long-horizon reasoning tasks. This limitation is particularly acute in multi-turn interactive reasoning, where learning is hindered by sparse terminal rewards, ambiguous trajectory-level credit assignment, and a compounding state-distribution mismatch between fixed replay trajectories and the evolving policy. In this paper, we propose HBR-GRPO, or Heuristic-guided Branching-Replay Group Relative Policy Optimization, a natural and effective extension of GRPO designed to be broadly applicable to multi-turn interactive reasoning. We adopt a heuristic-search perspective that decomposes trajectory-level training into individual turns and evaluates each action by the reduction it induces in the heuristic distance from the current state to the answer. We then use on-policy branching replay to reconnect these turn-level examples through transitions generated by the current policy, yielding genuine multi-turn interaction trajectories. HBR-GRPO substantially improves the performance of Qwen3-8B across the evaluated tasks. Notably, on one dataset where nearly all models with approximately 8B parameters listed on the clembench leaderboard achieve 0% accuracy, our method attains a nontrivial success rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.