acceptodds
Under review as a conference paper at ICLR 2027

Learning Safe Policies from Hard Counterfactual Negatives

Abstract

Safe reinforcement learning typically handles unsafe behavior by penalizing itduring training or preventing it from being executed. However, when an unsafe ac-tion offers competitive task value, these mechanisms do not explicitly require thepolicy to learn a safe alternative that is better. We introduce Hard-Negative Pol-icy Optimization (HNPO), a world-model-based method that uses such reward-competitive, near-boundary violations, or tempting failures, as hard negatives forpolicy learning. HNPO first identifies unsafe actions that remain credible un-der model uncertainty and behavioral support. A conservative hard-negative pro-poser then maintains access to informative negatives as the task policy becomessafer, learning from validated examples rather than directly optimizing throughworld-model gradients. To learn from these negatives, HNPO evaluates safe pol-icy actions and unsafe alternatives from the same latent state under a matchedcontinuation policy. Hard-Negative Policy Contrast constructs a utility referencefrom the validated negative group and optimizes safe policy actions to outperformthis reference by a margin, without treating proposer actions as on-policy sam-ples. HNPO leaves the backbone safe-RL objective and deployment-time safetymechanism unchanged. Across four Safety-Gymnasium continuous-control tasks,HNPO improves task reward over its backbone while maintaining low episodiccost, with larger gains on tasks exhibiting stronger reward–safety conflict. Abla-tions further demonstrate the complementary roles of credible negative selection,proposer-based coverage, and the negative-only reference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.