acceptodds
Under review as a conference paper at ICLR 2027

Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization

Abstract

To encourage diverse exploration in reinforcement learning (RL) for large language models (LLMs) without compromising accuracy, we propose Policy Split, a paradigm that bifurcates the policy into normal and high-entropy modes with a high-entropy prompt. While sharing model parameters, the two modes undergo collaborative dual-mode entropy regularization tailored to distinct objectives. Specifically, the normal mode optimizes for task correctness, while the high-entropy mode incorporates a preference for exploration, and the two modes learn collaboratively. Experiments across model sizes show consistent improvements in average reasoning accuracy together with entropy recovery. On creative-writing tasks, the trained high-entropy mode also produces more diverse outputs and improves several judged quality dimensions. Further analysis shows that the two modes develop distinct behavioral patterns and provide complementary learning signals. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.