acceptodds
Under review as a conference paper at ICLR 2027

Safety Alignment with Adaptive Preference Weighting

Abstract

Safety alignment should discourage harmful assistance while keeping useful answers to benign requests. Preference pairs specify which response to prefer, but averaging their losses can hide differences in how harmful and benign preferences are learned. We introduce Adaptive Preference Weighting (APW), which tracks harmful and benign preference losses separately and adjusts their coefficients within bounded intervals. APW scores responses by their mean log probabilities relative to a frozen reference model and trains one adapter with the resulting weighted loss. When training reuses the model's current errors, APW also caps the share of each harmful category and clips example weights, so that no category or example dominates. With the same base model, 5,333 ordered training pairs, adapter budget, decoder, and evaluation pipeline, APW obtains an Overall error (the mean of the harmful and benign error rates) of 18.20%, lower than all seven preference-optimization baselines. It also has the lowest error on a 600-prompt extension subset, 20.29%. Its harmful error is lower than that of every baseline, and this reduction is larger than its increase in benign refusal. Source breakdowns and prompt-level comparisons show where APW corrects harmful answers. A separate study that trains on the model's current errors tests the replay controls, sampled decoding, and the choice of refusal detector. Tracking harmful and benign losses separately is thus a practical way to set how strongly each kind of preference shapes safety training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.