acceptodds
Under review as a conference paper at ICLR 2027

DePO: Mitigating Policy Entropy Collapse via Denoised Policy Optimization

Abstract

Reinforcement learning has become a central tool for improving large language model reasoning, but rapid policy entropy decay can restrict exploration during training. We propose Denoised Policy Optimization (DePO), which filters importance sampling ratios along token positions and uses the filtered values as detached weights in a policy update with sign-dependent clipping gates. We study sliding-window averaging, exponentially weighted averaging, and Savitzky–Golay filtering as simple ways to control local ratio fluctuations. Building on prior analyses of entropy dynamics, we derive the detached-weight gradient and characterize noise reduction and smoothing bias under explicit assumptions. Experiments on mathematical reasoning and code generation cover three model backbones. On Qwen3-4B, DePO improves the reported average math accuracy from 65.5 to 68.4 and code accuracy from 39.0 to 42.3, while retaining higher final policy entropy than GRPO. Window studies, generation statistics, and a multi-turn search evaluation complement filter ablations and clipping-relaxation stress tests. These results support ratio filtering as a practical stabilization method; its effects on individual reasoning decisions remain to be established.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.