MHPO: Modulated Hazard-aware Policy Optimization for Stable Reinforcement Learning
Abstract
Controlling importance ratios in reinforcement learning with verifiable rewards requires retaining useful learning signals while limiting the influence of extreme policy deviations. Hard clipping can eliminate direct gradients for informative tokens, while unbounded ratio weighting can overemphasize outliers. We propose Modulated Hazard-aware Policy Optimization (MHPO), which separates these concerns through its gradient multipliers. A Log-Fidelity Modulator (LFM) preserves local ratio fidelity and supplies a smooth, positive, globally bounded multiplier. A Decoupled Hazard Penalty (DHP) uses established cumulative hazard functions to attenuate ratio expansions and contractions separately, without reversing their direct gradient contributions. Under an explicit assumption on the joint moment of advantages and score norms, we derive a local bound on the second moment of the gradient. We also explain how repeated ratio contractions accumulate along autoregressive prefixes. Component ablations, token masking, and diagnostics of directional tails connect these design choices to retained learning signals and reduced prefix contraction. Among the evaluated baselines, MHPO achieves the best average across benchmarks on both Qwen3-4B-Base and Qwen3-30B-A3B-Base. Further evaluations on instruction-tuned and multimodal models support transfer across the tested settings
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.