acceptodds
Under review as a conference paper at ICLR 2027

Scaling Practical Circuit Discovery for Adversarial Attacks on Language Models

Abstract

Circuit level interpretability has developed almost entirely on templated tasks, which limits its applicability to real world settings. We scale it to safety classification and refusal over diverse, unstructured prompts by using adversarial suffixes as contrastive pairs, and search for circuits directly in the vast MLP neuron basis. Gradient scoring methods such as attribution patching and integrated gradients fail at this granularity because they score every neuron once, at the empty circuit, as if each acted alone. We instead re-score along the search trajectory, proposing block additions and removals from orderings evaluated at the current circuit and accepting only moves that causal evaluation confirms. This recovers circuits far smaller than existing baselines while satisfying standard circuit metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.