acceptodds
Under review as a conference paper at ICLR 2027

Sycophancy and Jailbreak Compliance Share a Causal Axis in Language Models

Abstract

Ablation and probing studies have identified linear directions in LLM residual streams that mediate behaviors. However, work in this field has largely treated these behaviors as isolated features. Empirical and methodological evidence motivate us to study behaviors relationally. We study three such behaviors (safety, sycophancy, and knowledge conflict) in Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct. We extract directions by difference-of-means, intervene on each during generation, and measure override on all three behaviors against norm-matched random directions. In both models, removing a sycophancy direction lowers jailbreak success, and at the final prompt token the sycophancy and safety directions steer each other's behavior (8 of 9 direction–behavior pairs in Llama, 7 of 9 in Qwen). These results suggest that sycophancy and jailbreak compliance share a causal axis, and that directions governing behavior can be studied relationally, as features. We include additional experiments using this relational overlap to create a white-box jailbreak without explicit harmful requests, raising attack success on unseen JailbreakBench requests from 0.38 to 0.82.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.