acceptodds
Under review as a conference paper at ICLR 2027

When Native Fit Misleads: Evaluating Attention Surrogates under Positional Interventions

Abstract

Attention surrogates offer interpretable approximations of attention mechanisms, but strong reconstruction of native attention may not imply faithful prediction under positional interventions. We study this gap using controlled position-ID shifts that modify the positional information of a sequence suffix within a single attention head while keeping tokens, causal masks, and upstream representations fixed. We compare two parameter-matched surrogate families for RoPE attention in Qwen3-1.7B-Base: a smooth positional-field surrogate (Cubic Field) and a learned-frequency rotary surrogate (Rotary Field). Although native-validation KL consistently selects Rotary Field, we identify reproducible heads on OpenWebText and WikiText-103 for which it is less faithful to positional interventions than both Cubic Field and a no-change baseline. The misranking persists in model-output evaluations, although the output-level differences are small. An exploratory extension to T5 suggests that this phenomenon is not limited to RoPE. These results demonstrate that native reconstruction fidelity alone does not guarantee positional-intervention fidelity and highlight the need to evaluate attention surrogates directly on the mechanistic responses they aim to explain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.