FSPO: Fast Adversarial Training against Jailbreak Attacks via Soft Prompt Optimization
Abstract
As large language models (LLMs) have achieved remarkable performance and are increasingly deployed across diverse domains, ensuring their reliability and safety has become critically important. Among various threats, jailbreak attacks, which manipulate model to elicit unsafe or prohibited outputs, have become a major challenge. Although various defensive frameworks have been proposed, most existing defenses are computationally expensive. Moreover, since current defenses do not explicitly account for adaptive attacks, they are easily circumvented by strong adaptive attacks. To address these issues, we propose *Fast Soft Prompt Optimization*, which effectively enhances robustness by relaxing discrete prompt into a continuous probability simplex, called a soft prompt, and adaptively optimizing it. The resulting soft prompt can be directly deployed as a continuous prefix within the model prior to release. We evaluate our methods under two information-rich settings: model-aware but defense-unaware static attacks, and model-and-defense-aware adaptive attacks. We further analyze attention shift mechanisms to provide insights into the underlying robustness improvements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.