Beyond the Final Token: Detecting Jailbreak Prompts via Chat-Suffix-Probed Safety Trajectories
Abstract
Prior works detect jailbreak prompts by extracting harmfulness and refusal signals from the hidden states at final tokens. Some jailbreaks, however, suppress these safety signals at the endpoint and appear benign to such detectors. We instead study how safety signals evolve during prompt processing. Our key insight is that the standard chat suffix can serve as a probe that elicits the model's safety assessment of the prefix observed so far. Appending the chat suffix to successive prompt prefixes can reveal latent refusal tendencies throughout the prompt. The rising and falling regions of the resulting refusal trajectory help locate harmful-content and jailbreak spans, respectively. Deleting identified jailbreak spans increases refusal scores and leads to safer responses, providing causal evidence for this interpretation. Building on these findings, we propose DualTrace, which tracks refusal and harmfulness trajectories, combines them into a four-channel sequence, and classifies the sequence with a temporal convolutional network. Across four 7B-scale chat models, DualTrace achieves 96.17–98.27% mean classification accuracy on jailbreak and benign prompts, outperforming five baseline methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.