acceptodds
Under review as a conference paper at ICLR 2027

MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning

Abstract

Controlling importance ratios in reinforcement learning with verifiable rewards requires retaining useful learning signals while limiting the influence of extreme policy deviations. Hard clipping can eliminate direct gradients for informative tokens, while unbounded ratio weighting can overemphasize outliers. We propose Modulated Hazard-aware Policy Optimization (MHPO), which separates these concerns through its gradient multipliers. A Log-Fidelity Modulator (LFM) preserves local ratio fidelity and supplies a smooth, positive, globally bounded multiplier. A Decoupled Hazard Penalty (DHP) uses established cumulative hazard functions to attenuate ratio expansions and contractions separately, without reversing their direct gradient contributions. Under an explicit assumption on the joint moment of advantages and score norms, we derive a local bound on the second moment of the gradient. We also explain how repeated ratio contractions accumulate along autoregressive prefixes. Component ablations, token masking, and diagnostics of directional tails connect these design choices to retained learning signals and reduced prefix contraction. Among the evaluated baselines, MHPO achieves the best average across benchmarks on both Qwen3-4B-Base and Qwen3-30B-A3B-Base. Further evaluations on instruction-tuned and multimodal models support transfer across the tested settings

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.