acceptodds
Under review as a conference paper at ICLR 2027

Seed-Level Instability in LLM Safety Alignment: Regime Shifts in Large-Margin Preference Optimization

Abstract

As large language models (LLMs) are deployed more widely, reliable safety alignment becomes increasingly important. However, we identify seed-level safety-regime instability in large-margin, label-based SafeDPO: different training seeds produce sharply different refusal profiles. Across 20 seeds on Qwen2.5-7B-Instruct, OR-Toxic prefix-refusal rates range from 1.83% to 98.47% at a fixed endpoint, while instruction-following accuracy remains comparatively stable. Cross-seed averages obscure these differences in refusal behavior. We introduce Support-Completion SafeDPO (SC-SafeDPO), which restores discarded both-unsafe comparisons using their annotated relative-safety ordering, zero additional margin, and a fixed positive weight. The method preserves the original main-comparison orderings and margins and requires no additional inference-time computation. A local affine analysis relates comparison coverage to seed sensitivity. Budget-matched ablations indicate that both mixture reweighting and restored preference gradients contribute to stability. In the paired 20-seed comparison, SC-SafeDPO reduces cross-seed refusal-rate variance by approximately 77% on XSTest unsafe prompts and 87% on OR-Toxic relative to Label SafeDPO. It also lowers mean HarmBench attack success and benign false-refusal rates and improves instruction-following accuracy. Evaluations across three model backbones examine safety–helpfulness trade-offs and cross-seed variability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.