Fine-grained Credit Assignment for Multi-turn Critic-free LLM Reinforcement Learning with Sequential Equilibrium Optimization
Abstract
Critic-free group-based methods have shown strong potential for single-turn large language model (LLM) reinforcement learning (RL), owing to their reduced memory requirements and stable optimization. However, complex tasks often require multi-turn interactions, where the sparsity and delayed nature of verifiable rewards pose a fundamental challenge for credit assignment. Although recent methods alleviate this issue by introducing additional step-level groups or related mechanisms, trajectory-level grouping remains an essential component of advantage estimation and thus continues to exert a substantial influence on optimization. This motivates us to ask whether credit assignment can be further improved from a complementary perspective by directly refining the computation of trajectory-level advantages. To this end, we propose Sequential Equilibrium Policy Optimization (SEPO), a general credit assignment algorithm that can be flexibly integrated with a variety of existing group-based methods. SEPO introduces a novel multi-agent perspective that formulates multi-turn interaction as a fully cooperative sequential game. It then derives a solution to the subgame perfect Nash equilibrium through best-response optimization and aligns this solution procedure with multi-turn LLM RL, yielding more accurate trajectory-level credit assignment. Extensive experiments on ALFWorld and WebShop with Qwen2.5-1.5B, Qwen2.5-7B, and Qwen3-4B demonstrate that SEPO consistently improves over GRPO, GiGPO, and DAPO while holding the amount of data used for actor updates unchanged, validating the effectiveness and generality of its credit assignment mechanism.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.