acceptodds
Under review as a conference paper at ICLR 2027

FSPO: Fast Adversarial Training against Jailbreak Attacks via Soft Prompt Optimization

Abstract

As large language models (LLMs) have achieved remarkable performance and are increasingly deployed across diverse domains, ensuring their reliability and safety has become critically important. Among various threats, jailbreak attacks, which manipulate model to elicit unsafe or prohibited outputs, have become a major challenge. Although various defensive frameworks have been proposed, most existing defenses are computationally expensive. Moreover, since current defenses do not explicitly account for adaptive attacks, they are easily circumvented by strong adaptive attacks. To address these issues, we propose *Fast Soft Prompt Optimization*, which effectively enhances robustness by relaxing discrete prompt into a continuous probability simplex, called a soft prompt, and adaptively optimizing it. The resulting soft prompt can be directly deployed as a continuous prefix within the model prior to release. We evaluate our methods under two information-rich settings: model-aware but defense-unaware static attacks, and model-and-defense-aware adaptive attacks. We further analyze attention shift mechanisms to provide insights into the underlying robustness improvements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.