From Embeddings to Attention Heads: Input-Agnostic Tracing of Embedding-MLP-Attention Circuits in LLMs
Abstract
Input-agnostic mechanistic interpretability explains neural computation from model parameters without repeated forward passes. Existing input-agnostic methods still face a tradeoff: analyses of interactions between modules are mostly limited to toy transformers, while studies of standard LLMs usually focus on isolated modules, leaving key residual-stream interactions underexplored. We address this gap by introducing a task-conditioned zero-pass framework for tracing embeddingnative gated-MLP-neuronattention-head interactions in standard LLMs. Our method requires no task-example forward passes for activation/gradient collection. To our knowledge, prior input-agnostic work has not traced these interactions at this scale in a standard pretrained LLM. Experiments against input-dependent methods, including state-of-the-art causal path patching methods, show that our method achieves a highly favorable balance between causal effectiveness and computational efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.