When Native Fit Misleads: Evaluating Attention Surrogates under Positional Interventions
Abstract
Attention surrogates offer interpretable approximations of attention mechanisms, but strong reconstruction of native attention may not imply faithful prediction under positional interventions. We study this gap using controlled position-ID shifts that modify the positional information of a sequence suffix within a single attention head while keeping tokens, causal masks, and upstream representations fixed. We compare two parameter-matched surrogate families for RoPE attention in Qwen3-1.7B-Base: a smooth positional-field surrogate (Cubic Field) and a learned-frequency rotary surrogate (Rotary Field). Although native-validation KL consistently selects Rotary Field, we identify reproducible heads on OpenWebText and WikiText-103 for which it is less faithful to positional interventions than both Cubic Field and a no-change baseline. The misranking persists in model-output evaluations, although the output-level differences are small. An exploratory extension to T5 suggests that this phenomenon is not limited to RoPE. These results demonstrate that native reconstruction fidelity alone does not guarantee positional-intervention fidelity and highlight the need to evaluate attention surrogates directly on the mechanistic responses they aim to explain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.