acceptodds
Under review as a conference paper at ICLR 2027

Form Before Harm: Diagnosing the Semantic Influence of Jailbreak Formats with Activation Steering

Abstract

Large language model safety defenses often organize harmful inputs by topic without accounting for the semantic shifts induced by jailbreak formats. We show that these shifts can alter this organization: attack format has a substantially greater impact on the distribution of LLM semantic representations than harmful topic. To study this effect in activation space, we introduce Context-Aware Steering (COS), an adaptive activation-steering method that learns a routable steering space from mixed jailbreak data and achieves strong defense performance with limited benign-side effects. Across both the model's natural representations and the intervention structure learned by COS, attack format consistently plays a larger organizing role than harmful topic; COS further adapts its shared and specialized steering components to the internal structure of each LLM rather than relying on predefined topic categories. We further identify a clear out-of-distribution generalization gap across jailbreak formats. Increasing the diversity of attack formats during training alleviates this gap but does not eliminate it. Together, these findings show that jailbreak format is a major determinant of LLM semantic representations and should be considered when organizing safety detection and defense.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.