acceptodds
Under review as a conference paper at ICLR 2027

Attention Tracks Reaction Centers, But Doesn't Need Them: A Causal Study of a Retrosynthesis Transformer

Abstract

Attention weights in transformer-based chemistry models are widely treated as an interpretability signal — attention concentrated on a molecule's reaction center is commonly cited as evidence that a model has "learned" the relevant chemistry. We test this assumption directly using causal intervention. In a retrosynthesis Chemformer trained on USPTO-50k, encoder attention correlates strongly with ground-truth reaction centers (within-molecule AUC 0.902 vs. 0.836 for hand-engineered structural descriptors), replicating the correlational finding that motivates this interpretability practice. However, zero-ablating attention heads and measuring the resulting change in generation likelihood reveals that causal necessity is inversely related to this correlation: the heads most predictive of reaction centers are among the least necessary for generation, while 74.9% of total ablation damage concentrates in the encoder's first layer — the layer with the lowest reaction-center correlation of any layer in the network. We show this necessity is not decomposable into individual head contributions: the sum of single-head ablation effects in layer 1 is 115× smaller than the effect of ablating all eight heads jointly, and which single head survives a near-total ablation determines whether the model recovers 84% or 31% of baseline performance. We trace this structure to a measurable property of each head: attention entropy predicts sole-survivor sufficiency (ρ = −0.857) in the opposite direction from necessity, explaining why sharp, low-entropy positional heads are individually redundant yet individually sufficient, while diffuse, high-entropy heads are collectively necessary yet cannot reconstruct function alone. Our results extend the attention-is-not-explanation literature from NLP classification into generative chemistry models, and show that correlational interpretability claims in this domain — as in concurrent findings on 3D-geometry alignment in other chemical language models — do not reliably predict where causal computation actually occurs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.