Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Abstract
Inoculation prompting (IP) is a selective-generalization technique used against emergent misalignment. We study inoculation adapters (IA), a family of methods that similarly reduce the optimization pressure to learn undesired traits by strengthening those traits during training. IAs are LoRAs used in three steps: (1) trained on undesired traits; (2) attached frozen while a separate task adapter is trained on data exhibiting both desired and undesired traits; (3) the IA is discarded at deployment, while only the task adapter is kept. Vanilla IA shares its mechanism with security vectors (Zhou et al., 2024); we introduce two gated variants. We compare IAs with four selective-generalization baselines: IP, preventative steering, Concept Ablation Fine-Tuning (CAFT), and KL regularization. Across nine setups and five model families, the IA family spans the observed Pareto frontier of desired trait retention vs. undesired trait suppression among the evaluated methods, although given wide confidence intervals the magnitude of improvement remains uncertain. IAs also avoid two drawbacks of IP: they can suppress capabilities and traits that cannot be reliably elicited by a prompt, and they introduce fewer surprising backdoors. However, no IA variant optimizes all objectives perfectly; gains in desired-trait generalization are generally accompanied by weaker suppression of the undesired trait and increased backdoor occurrence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.