acceptodds
Under review as a conference paper at ICLR 2027

Refusal is non-stationary: Iterative steering for LLM uncensoring

Abstract

Despite recent progress in safety alignment, Large Language Models (LLMs) remain vulnerable to steering attacks that bypass safety mechanisms and produce uncensored models that no longer refuse harmful requests. Existing methods typically estimate fixed refusal-related directions and apply them through one-shot interventions, yielding limited success on strongly aligned models or substantially degrading performance on benign tasks. We formulate uncensoring as a constrained steering problem that minimizes refusal on harmful prompts while limiting behavioral deviation on benign inputs, thereby discouraging interventions that reduce refusal through broad corruption of model behavior. We further show that refusal is non-stationary under steering: each intervention changes how residual refusal is represented, causing previously estimated directions to become progressively less effective. Motivated by this observation, our method iteratively re-estimates refusal after each intervention, recomputes layer-specific steering directions, and selects successive interventions subject to the benign-behavior constraint. Across eleven LLMs, our approach consistently achieves higher attack success rates than existing steering methods, with particularly large gains on larger and more strongly aligned models where prior approaches largely fail. Finally, the ordered sequence of interventions produced by our method provides a tool for mechanistic interpretability, revealing how refusal is distributed across layers and how its representation evolves throughout the steering process.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.