acceptodds
Under review as a conference paper at ICLR 2027

SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

Abstract

Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model’s refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: (1) stronger alignment, lowering harmfulness score by 63.21%; (2) less over-refusal, yielding a 58.44% decrease in refusal rates for benign queries; and (3) better utility, retaining 99.58% of the original model capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.