Jailbreaks Live in Persona Space
Abstract
Understanding why aligned language models fail is central to AI safety. Jailbreaks provide a concrete instance, where adversarial prompts can shift models from refusal to harmful compliance. We study jailbreaking through the emerging view of LLMs as persona simulators, asking whether jailbreak behavior is organized within persona space. On Llama-3.1-8B and Gemma-4-31B, we find that jailbreaks are strongly structured by persona space. Individual traits predict jailbreak outcome, persona structure emerges in principal components learned solely from jailbreak activations, and persona-space projections retain most of the predictive signal in refusal and jailbreak directions. Building on this structure, we construct the PersonaJB Direction, a single direction constrained entirely to persona space that achieves the strongest average generalization to unseen attacks on both models with minimal inference overhead. Finally, steering along PersonaJB and individual trait directions induces jailbreaks at rates comparable to specialized safety directions and standard attacks. Our results establish persona space as an interpretable, predictive, and causal representation of jailbreak behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.