Path-Space: Estimating Transformer Component Effects and Their Interactions from Shared Backward Passes
Abstract
Which components of a transformer matter for a prediction, and which matter only in combination? Ablation answers this at one forward pass per component and one per pair, and the cheap readings used instead, the logit lens and direct-logit attribution, ignore everything downstream of the component. We present Path-Space, a method that estimates component effects and their pairwise interactions from shared backward passes. It reads each block’s write through the exact per-context transport to the output, as attribution patching does, at every block and position from one backward pass; it obtains a removal’s exact effect as the line integral of the gradient along the removal, evaluated by Gauss–Legendre quadrature; and it introduces second-order attribution patching, which predicts a component’s interaction with every other component from one backward pass on the run with that component removed, with an exact error identity and a symmetrised form. On five models from 0.36B to 70B parameters, scored against measured removals on the model’s own next-token log-probability, the per-context reading predicts removal effects at Spearman 0.69–0.82, against 0.20–0.29 for the identity reading behind the logit lens; six quadrature nodes reproduce the effects to a median error below 2 × 10−4 at every block but the first; and the symmetrised interaction estimate reaches 0.65–0.76 up to 8B and 0.56 at 70B in bf16, recovering strong interactions between individually weak components that rankings by single effects miss. Its setup costs about 4L forward passes against L 2/2 for all pairs: at 80 blocks it beats measuring pairs in the order of their single effects at equal cost from a fifth of all pairs on, a comparison we did not pre-register, while at 24–32 blocks it does not yet pay for itself
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.