acceptodds
Under review as a conference paper at ICLR 2027

HIPO: Heterogeneous and Instantaneous Policy Optimization for Safe Language Model Alignment

Abstract

Safety alignment of large language models seeks to maximize helpfulness while satisfying safety constraints. Existing approaches typically impose safety pressure through global or stage-wise control, which can overlook that safety violations are heterogeneous across prompts and evolve throughout training. We propose HIPO (*Heterogeneous and Instantaneous Policy Optimization*), a two-level safety-control mechanism driven by the model's current responses. At the *prompt level*, current positive violations determine heterogeneous safety weights, assigning stronger correction to riskier response groups. At the *rollout level*, the aggregate safety pressure calibrates the learning rate, enabling stronger safety-directed updates while keeping the update stable. We establish HIPO achieves bounds on cumulative positive reward regret and constraint violation with accurate exact policy gradient. Experiments on Alpaca-7B and Qwen3.5-9B across four safety benchmarks show that HIPO improves safety while preserving strong helpfulness, as justified by reward–cost evaluation, pairwise LLM judgments, and benchmark-specific safety metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.