A Faithful Attribution Reports Its Head, Not Its Model: The Limits of Completeness in Additively Decomposed Representations
Abstract
Additive models promise an explanation readable off the architecture: the latent is a sum of per-component terms, so each contribution is available without a post-hoc estimator. That promise is bounded by the downstream head. A completeness-satisfying attribution decomposes a head's output, so what it can say is fixed by the head you chose, not by the model. On MD-TJEPA, a JEPA whose latent decomposes over channel band scale, channel identity is decodable from the components and only (chance ) from the sum they form. Fit the head for the channel and its exact decomposition closes of the gap to a pre-aggregation score (, ); on a plain NAM the same construction is not weak but arbitrary, or in most seeds. Fit it for the model's own task — a band detector whose labels contain no channel — and an equally exact decomposition names the channel at . An input intervention separates the two: resampling the injected channel moves both heads' outputs, moving its content to another channel barely does, so the sum hides which channel contributed, not whether it did. The band head's map agrees with that intervention on of windows, the channel head's on ; most of each map's disagreement is its implicit reference, the latent origin. Re-referenced to the training-mean contribution — on this synthetic benchmark an estimate of the expected intervention — they agree on and . Completeness pins the sum, not the head and not the reference. We are scoped on faithfulness, not accuracy — TS2Vec outperforms MD-TJEPA under a matched probe, though an untrained MD-TJEPA read through its components outperforms TS2Vec — and the axes conflated as "interpretability" disagree on the same checkpoints. What this buys is a check worth running first: probe the aggregate for the axis, and intervene on the input for it, before trusting any attribution over it — a probe near chance says the aggregate hides which channel was active, not that its content goes unused. between and costs one line and exposed a gate collapse here — the JEPA objective shares the gate between prediction and target, so an unopposed sparsity penalty switches components off. In a minimal additive stand-in, removing the gate from the target prevents that shrinkage outright. We close with three checks to run before reading any additive attribution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.