acceptodds
Under review as a conference paper at ICLR 2027

Recruited and Erased: Head-Level Provenance of Safety Detection in Language Models

Abstract

Aligned language models refuse harmful requests, yet the dominant mechanistic account collapses harm detection and refusal execution onto one residual-stream direction. It cannot say which heads detect harm, or whether safety fine-tuning built or found them. We introduce a Safety Specificity Score () that selects heads whose output shifts more on harmful than benign prompts. To trace their origin, we score each selected head at three checkpoints, base, instruct, and abliterated, whose refusal direction is removed, across six checkpoint sets. Fine-tuning raises on the heads abliteration pushes back down. A Holm-corrected Wilcoxon places instruct above abliterated on Qwen3-8B and Gemma-3-4B, and the direction replicates on two of four further sets. On Qwen3-8B, fine-tuning recruited 16 of the 19 heads abliteration erases. Detection heads are recruited, not built: harm-matched, base checkpoints already reach 0.74 to 0.80 AUROC through the instruct-trained probe, below the 0.852 of an input-only classifier, so what the recruited heads read out is bounded by the prompt text. Zeroing them never reproduces abliteration's refusal collapse: detection is separable from execution, and head-level provenance measures a stage that direction-editing defenses act beside rather than measure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.