WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior
Abstract
Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often require task-specific training, lack natural language controllability, or compromise semantic coherence.To bridge this gap, we propose WASD (unWeaving Actionable Sufficient Directives), a mechanistic interpretability framework that explains model behavior by identifying intervention-validated sufficient neural conditions for token generation. WASD uses Circuit Tracer as an attribution and activation extractor, but does not directly treat high-attribution neurons as explanations. Instead, it constructs neuron-activation predicates under prompt perturbations and searches for a compact, pruned rule set whose intervention preserves the target behavior. Experiments on SST-2, CounterFact, and IOI across four models yield 94.83%–99.48% mean fidelity for compact rules at a target threshold of 0.95; protocol differences limit direct comparisons with attribution baselines. Neutral-prefix evaluations at the same threshold show lower mean instability in all twelve settings. Cross-lingual completion and QA on Gemma-2-2B and Qwen3-8B demonstrate language control and task-dependent trade-offs between accuracy and Chinese explanation rate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.