acceptodds
Under review as a conference paper at ICLR 2027

Why Safety Alignment Can Fail: A Geometric Theory of Structural Hallucination in LLMs

Abstract

Safety alignment can fail under jailbreak even measured task capability remains intact, but a change in a safety projection alone does not reveal whether the representation moves farther or moves in a different direction. We give an exact geometric decomposition of every normalized inter-layer update into direction-independent retention and a direction-specific tangent contribution, derive the sharp reachable interval of a single update, and prove a finite-depth directional-budget bound with a closed-form characterization of equality. Exponential projection decay follows as the constant-retention, zero-budget special case. Across four Llama and Qwen models on HarmBench, refused trajectories receive systematically larger safety-directed tangent contributions than jailbroken trajectories, with refused-minus-jailbroken effect sizes from 1.3 to 2.3. The retention contrast has the opposite sign and partly offsets the tangent contrast. The tangent ordering remains after adjustment for layer and current projection. A separate Llama-3.1-8B GSM8K experiment observes no statistically distinguishable accuracy difference between clean prompting and the tested jailbreak prefix. These results support analyzing safety maintenance through both retention and direction-specific writes, clarifying how task computation can persist while constraint transmission weakens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.