acceptodds
Under review as a conference paper at ICLR 2027

Ordered Advantages Prevent Safety Drift in RL Fine-Tuning

Abstract

Which signals in a model's rollouts become persistent behavior during reinforcement-learning (RL) fine-tuning? We show that group-relative advantages determine this channel: FP8 rollouts increase rollout–trainer mismatch by 2.8–7.7× while safety remains in the BF16 range across four seeds, whereas a reward that favors harmful responses raises Qwen2.5-7B-Instruct harmful compliance from 6.7% to 87.8% in 100 steps. We introduce Ordered Advantage Projection (OAP), a provider-side method that projects each response group so a more harmful response never receives a larger advantage. Isotonic regression computes the projection, which preserves the group mean, requires no threshold or penalty weight, and is invariant to strictly increasing score calibration. In the default harmful-reward attack, OAP with a KL anchor gives 6.9% AdvBench harmful compliance versus 6.7% for the base model and 87.8% without defense. Across four seeds, five adaptive attacks, two provider classifiers, two RL estimators, and three models, OAP retains near-base safety and helpfulness; its projection adds 0.28% to step time. Under a benign GSM8K reward, OAP and no defense reach 78.1% accuracy, while penalty baselines reach 77.3–77.7%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.