Learning Safe Policies from Hard Counterfactual Negatives
Abstract
Safe reinforcement learning typically handles unsafe behavior by penalizing itduring training or preventing it from being executed. However, when an unsafe ac-tion offers competitive task value, these mechanisms do not explicitly require thepolicy to learn a safe alternative that is better. We introduce Hard-Negative Pol-icy Optimization (HNPO), a world-model-based method that uses such reward-competitive, near-boundary violations, or tempting failures, as hard negatives forpolicy learning. HNPO first identifies unsafe actions that remain credible un-der model uncertainty and behavioral support. A conservative hard-negative pro-poser then maintains access to informative negatives as the task policy becomessafer, learning from validated examples rather than directly optimizing throughworld-model gradients. To learn from these negatives, HNPO evaluates safe pol-icy actions and unsafe alternatives from the same latent state under a matchedcontinuation policy. Hard-Negative Policy Contrast constructs a utility referencefrom the validated negative group and optimizes safe policy actions to outperformthis reference by a margin, without treating proposer actions as on-policy sam-ples. HNPO leaves the backbone safe-RL objective and deployment-time safetymechanism unchanged. Across four Safety-Gymnasium continuous-control tasks,HNPO improves task reward over its backbone while maintaining low episodiccost, with larger gains on tasks exhibiting stronger reward–safety conflict. Abla-tions further demonstrate the complementary roles of credible negative selection,proposer-based coverage, and the negative-only reference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.