acceptodds
Under review as a conference paper at ICLR 2027

Reading Multiplicative Circuits in Transformer MLPs

Abstract

Transformer MLPs are usually read through activations, sparse feature dictionaries, or logit-space projections. These views underuse the multiplicative structure of gated feed-forward blocks, where the native unit of computation is an interaction between two sources. Circuit methods built on activation patching estimate a patch’s effect with a gradient, which is linear in the patch, so whatever two sources contribute jointly is assigned zero. We measure that joint contribution with a 2 × 2 ablation and find it runs from 0.12 to 1.19 times the size of the main effects those methods recover accurately. We then introduce behavior-conditioned integrated-Hessian tomography, a weight-native method that reads which residual-stream sources interact inside a pretrained gated MLP. Decomposing the normalized MLP input into named upstream sources decomposes the output into pairwise source interactions, and the decomposition is exact: the terms sum back to the MLP’s true output to within one part in ten million. It predicts the measured interaction at Pearson correlation 0.95 to 0.97 over 1152 pairs at four layers. The same construction applies to attention scores, which are already exactly bilinear and need no path integral, so the method covers two of the three computational units of a transformer block. Patching a pair’s write moves the next-token log-probability by the predicted amount, at slope 0.98 with R2 = 0.91. The terms that compose the MLP output are about three orders of magnitude larger than the output itself, with output cancellation growing as the square of input cancellation, so ranking pairs by the magnitude of their write recovers which sources are large. Projecting each write onto the layer output gives an exact signed attribution that accounts for the cancellation and ranks pairs independently of their size.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.