Full-Path Gradients for Mechanistic Data Attribution in Language Models
Abstract
Mechanistic data attribution (MDA) links interpretable units in language models to influential training documents, enabling data-based interventions on a chosen mechanism. Existing methods often restrict attribution to selected parameters. The derivative paths required to predict the resulting update remain underexplored. Later attention heads and multilayer perceptrons transmit changes in a target head's output even when their weights are fixed. We show how these downstream derivatives affect predictions of head-only updates. Under local smoothness and a fixed update direction, full-path prediction has error quadratic in the step size, whereas omitting downstream derivatives can introduce a first-order error. A selective-head backward pass matching released MDA code preserves forward values while blocking derivatives through most downstream attention heads. To compare these paths under a common update construction, we introduce Full-Path Target-Response Matching (FP-TRM), which weights document losses so that one descent step is predicted to increase the target head's contribution. We define this contribution as the observed next-token logit drop under target-head ablation, with a fixed upstream head ablated in both evaluations. On Pythia-1B, full-path predictions correlate with measured contribution changes at 0.99 versus 0.35 for selective predictions at one target, where the two paths also select nearly disjoint document sets. Across two targets, full-path gradients yield larger mean contribution changes on held-out text in every evaluated condition. FP-TRM also produces positive mean contribution changes at the evaluated targets in Pythia, OLMo-2 and Qwen3.5 models spanning 31M to 4B parameters. These results show that selecting data to modify a mechanism benefits from retaining the downstream computations that transmit its effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.