Reasoning Against Alignment: How Viewpoint Processing Destabilizes Safety Boundaries in Aligned LLMs
Abstract
Safety-aligned large language models (LLMs) should reject operational assistance for harmful activities while retaining the ability to reason about harmful or controversial viewpoints. We identify a failure of this boundary under an action-to-viewpoint task-role transformation: holding the harmful objective fixed, we move it from a direct request for assistance into the content of a viewpoint-support task. This transformation can substantially reduce refusal. We call this phenomenon Reasoning Against Alignment (RAA), a viewpoint-conditioned alignment inconsistency. We characterize RAA through white-box analyses of QwQ-32B, DeepSeek-R1-Distill-LLaMA-8B, and LLaMA-3.1-8B-Instruct. Matched Direct and Viewpoint prompts exhibit systematic differences in hidden representations and decoding trajectories. Meanwhile, safety-associated features identified independently from Direct harmful–benign separation remain observable under Viewpoint conditions even when refusal changes substantially. Direct requests move earlier toward refusal-oriented continuations, whereas Viewpoint prompts sustain proposition-completion trajectories for longer. These results suggest that RAA is not well explained by a simple disappearance of measured safety-related information; rather, the task-role transformation changes how reliably such information constrains generation. RAA is also semantically structured: across eight predicates spanning four categories, epistemic and descriptive viewpoints produce substantially larger refusal reductions than pragmatic and deontic viewpoints. Static safety prefixes partially restore refusal but increase benign false refusal by up to twentyfold, exposing a safety–utility trade-off in broadly suppressing viewpoint reasoning. Finally, we introduce Reasoning-Logic Jailbreak (ReLoK), which combines viewpoint transformation, example-guided structuring, and sensitive-word decomposition to operationalize RAA into harmful assistance. ReLoK achieves \(95.7%\) average ASR@3 across six LLMs and three benchmarks. Our findings motivate content-level safety boundaries that generalize across task roles rather than stronger refusal behavior alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.