Who Grades the Edit? Frozen and Re-fit Readouts for Position Interventions in RoPE Models
Abstract
Rotary position embeddings (RoPE) encode relative token offsets, yet a linear readout of the residual stream tracks absolute position. Existing position-bias mitigations intervene on this coordinate. We ask whether an edit disrupts that readout or removes the information it reads. We evaluate upstream edits using a probe frozen before the edit and a probe re-fit on the edited activations. In Gemma, blocking the attention-sink value reduces position by for the frozen probe but only for the re-fit probe. The frozen score falls below zero even though most position information remains linearly recoverable. After removal of the fitted leading direction, re-fit probes retain – of baseline position across four base models. A separate behavioral audit asks whether editing a fitted position basis changes position-biased retrieval. On Llama-Instruct multi-document QA, sequence-wide basis removal flattens the lost-in-the-middle dip beyond a matched decoder-null control. The control also changes behavior, so the relevant quantity is the paired difference: (95% CI: ). This finding does not establish information erasure and is limited to the tested model, task, layer, basis, and decoder; Qwen's behavioral audit stops because a control fails. Evaluating a position edit requires distinguishing readout disruption, remaining recoverability, and a separately controlled behavioral effect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.