Beyond Task-Conditioned Importance: Head-Resolved Output Jacobians Separate Broad Response from Behavioral Relevance
Abstract
Interpretability in explainable AI is often quantified by the importance of an internal component by how much it 'contributes' to a given prediction/task/behavior. Here we ask the orthogonal question: can an unimportant component for one behavior be highly gainful in other directions of output space? We formulate a head-resolved framework that disentangles gross downstream responsivity, task-guided sensitivity/alignment, and behavioral visibility at the prediction head. This is one of the first systematic comparison of head-output-to-vocabulary Jacobian response with both task-directed sensitivity and behavioral readouts across multiple language models. We study GPT-2 Small, Pythia-410M, and GPT-Neo-125M using output Jacobians, activation patching, and direct logit attribution. The clearest case occurs in Pythia-410M: its two most broadly responsive heads rank only 263rd and 374th in behavioral visibility on indirect object identification. Finite interventions closely match the Jacobian-predicted slopes (approximately 1% relative error at the tested scales), while normalized output gain remains comparable to high-response controls matched with sensitivity. We find that their finite-head gain also varies when we change the task contrasts. This demonstrates that poor importance of task-conditioned task need not correlate with weak local responses downstream. Gross gain, task-guided alignment, and behavioral visibility are therefore orthogonal views of a given model component, rather than synonyms for importance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.