PROOF: Mixed-Traffic Jailbreak Defense via Profile Routing in Frozen LLMs
Abstract
Safety controls must prevent harmful completions while preserving legitimate assistance. We introduce , which learns to select among fixed safety profiles of a frozen language model. During training, each input is evaluated under every profile using a cost appropriate to its type: unsafe completion, unnecessary refusal, or capability-probe loss. When several profiles share the lowest cost, a tie-aware target retains a preferred action and assigns probability to the other minima. This preserves equally rated alternatives that hard-label imitation would penalize. We study both routing before response generation and selection among profile-generated candidates. On a fixed 200-example outcome table, ten tie-breaking assignments and five paired controller seeds yield a 7.69% reduction in sensitivity to tie-breaking relative to hard-label training, with no statistically resolved change in mean selection regret. Across eight profile-menu conditions, safety differences depend on the available actions. Serving measurements further show that candidate conditioning costs more than routing from prompt features. The results identify a benefit of tie-aware supervision for label sensitivity and distinguish it from menu-dependent safety gains and inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.