acceptodds
Under review as a conference paper at ICLR 2027

Jailbreaks Live in Persona Space

Abstract

Understanding why aligned language models fail is central to AI safety. Jailbreaks provide a concrete instance, where adversarial prompts can shift models from refusal to harmful compliance. We study jailbreaking through the emerging view of LLMs as persona simulators, asking whether jailbreak behavior is organized within persona space. On Llama-3.1-8B and Gemma-4-31B, we find that jailbreaks are strongly structured by persona space. Individual traits predict jailbreak outcome, persona structure emerges in principal components learned solely from jailbreak activations, and persona-space projections retain most of the predictive signal in refusal and jailbreak directions. Building on this structure, we construct the PersonaJB Direction, a single direction constrained entirely to persona space that achieves the strongest average generalization to unseen attacks on both models with minimal inference overhead. Finally, steering along PersonaJB and individual trait directions induces jailbreaks at rates comparable to specialized safety directions and standard attacks. Our results establish persona space as an interpretable, predictive, and causal representation of jailbreak behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.