Boundary-Stable PolicyFlip: Training Policy-Sensitive Safety Guards
Abstract
Safety guards deployed under changing policies must sometimes reverse decisions when trusted state changes even if the underlying event does not. We introduce Boundary-Stable PolicyFlip, a fine-tuning objective for matched policy-state interventions. Writing a pair by its unsafe–safe score gap g and score center c, a strict threshold crossing requires g>2|c|: conventional pairwise ranking controls g but leaves c unconstrained. We therefore optimize the weaker-endpoint slack b=g/2-|c|=min(-d_safe,d_unsafe). At the default settings, its penalty is exactly worst-endpoint cross-entropy, yielding a threshold-aware aggregation of the same endpoint labels. BS-PF combines this boundary term with transition ranking; BG-PF additionally attenuates ranking as a pair moves inside the correct decision regions. Neither changes the guard architecture or inference path. On a controlled policy-intervention benchmark, BS-PF improves Qwen3.5-4B balanced accuracy over pointwise training by 5.25–6.50 percentage points across four raw-input evaluations. Against a ranking-plus-paired-CE control on Policy OOD, BS-PF improves transition direction by 4.08 points (95% interval [2.69, 5.47]). On a fresh machine-rendered holdout, BG-PF improves pair exactness by 2.30 points and transition direction by 5.25 points [3.04, 7.42] over Pair-CE. With unchanged loss constants on Llama-Guard-3-8B, it improves Policy-OOD BAcc over gap-only training by 7.46 points and pair exactness by 12.07 points. Across 12,434 public moderation records, aggregate macro BAcc remains within 0.03 points of Decision-only (85.27% versus 85.24%). Overall, threshold-aware pair optimization strengthens policy-transition tracking under a fixed training-step budget without adding inference-time machinery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.