acceptodds
Under review as a conference paper at ICLR 2027

WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior

Abstract

Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often require task-specific training, lack natural language controllability, or compromise semantic coherence.To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a mechanistic interpretability framework that explains model behavior by identifying intervention-validated sufficient neural conditions for token generation. WASD uses Circuit Tracer as an attribution and activation extractor, but does not directly treat high-attribution neurons as explanations. Instead, it constructs neuron-activation predicates under prompt perturbations and searches for a compact, pruned rule set whose intervention preserves the target behavior. Experiments on SST-2, CounterFact, and IOI across four models yield 94.83%–99.48% mean fidelity for compact rules at a target threshold of 0.95; protocol differences limit direct comparisons with attribution baselines. Neutral-prefix evaluations at the same threshold show lower mean instability in all twelve settings. Cross-lingual completion and QA on Gemma-2-2B and Qwen3-8B demonstrate language control and task-dependent trade-offs between accuracy and Chinese explanation rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.