acceptodds
Under review as a conference paper at ICLR 2027

Targeting Duplicate Token Heads: Localizing and Repairing a Safety Circuit in LLMs

Abstract

Large language models remain vulnerable to jailbreaks whose circuit-level basis is poorly understood. We introduce Adversarial Circuit Patching (ACP), a mechanistic framework for identifying causal contributors to safety-related computation in DPO-tuned 7B–8B LLMs (Llama, Mistral, Openchat). Applying ACP, we identify Duplicate Token Heads (DTHs), attention heads that bind repeated entity mentions within a prompt, as one such contributor. Perturbing a single token achieves up to 39.8% attack success, and patching just 5 of 1,024 heads causally recovers up to 46% of the resulting refusal failures, showing that task-specific circuit components such as DTHs can be identified and exploited to attack safety-aligned models. We construct a curated Wikidata Private Information (WPI) dataset that exposes both task-relevant and safety-relevant circuit components, and use it to show that this vulnerability varies sharply by architecture: Llama-DPO's refusal circuit collapses under minimal perturbation, while Mistral and Openchat remain comparatively robust. We further propose a deployable mitigation, a single fixed steering vector per DTH, that restores refusal on over 50% of previously successful attacks for Mistral and Openchat (up to 100%) but only 15.4-55.7% for the more attack-susceptible Llama-DPO. By localizing exploitable safety-relevant behavior to a small, targetable circuit and demonstrating a lightweight repair for it, ACP moves beyond surface-level alignment toward targeted, circuit-level defenses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.