SafeCompose: Branch-Preserving Safety Composition for Held-Out Task Adapters
Abstract
Parameter-efficient fine-tuning delivers new capability as a small adapter that leaves the base model untouched, so the party that trains an adapter is often not the party that deploys it. The deploying party, which we call the operator, receives finished adapter weights but nothing of the training run behind them. If that run included unsafe examples, composing the adapter with an aligned base model can erode the model's refusal of unsafe requests. We ask whether a defense can restore refusal without giving up the adapter's capability, even for adapters it has never seen. We introduce SafeCompose: the aligned base, the received adapter, and a safety adapter stay frozen and separate, and small per-layer gates read the input to set how much each adapter contributes. The gates are trained once on a pool of task adapters and reused unchanged on adapters that arrive later. We test on new adapters across model families and levels of unsafe training data, including a task the gates never saw, and also over a sequence of arriving adapters and under attack. On the unseen task, SafeCompose restores refusal and keeps more task capability than weight-projection and activation-steering defenses, retaining the 9–31 points of accuracy that discarding the adapter would give up. Over successive adapter arrivals, it stays safe where steering defenses do not. When the response is forced to begin by agreeing, even by an attacker who sees the gates, SafeCompose is safer than fixed gates trained the same way. It keeps this advantage even when its gates read the prompt only once, which allows cached decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.