acceptodds
Under review as a conference paper at ICLR 2027

The Limits of Directional Alignment as Evidence for a Shared Mechanism

Abstract

The directional alignment of residual stream directions in language models has been put forth as evidence of a shared underlying mechanism. Since directional alignment and causal transfer may both occur at the same time, it is unclear whether directional alignment alone is enough to support a shared underlying mechanism. This question is tackled in the context of self-preservation, using a role × evidence conflict contrast in lieu of self-preservation contrast, and hence the directions obtained cannot be attributed solely to self-relevance/self-preservation. The rank-1 directions of Qwen3-8B trained separately across mutually exclusive scenarios in English, Chinese, Spanish, and Hindi languages display high alignment of 0.76-0.90 signed cosine similarities, compared with 0.26-0.42 under label permutation null. The selective causal control test is carried out via ablation of directions at each depth level. Role-conditioned decision gap falls within the label permutation null distribution in all but one language × depth condition. The latter also violates specificity the most. The intervention procedure causes a reduction of 0.844 in refusal on held out prompts as a positive control, indicating the capability of the intervention to influence model behavior. Further limitations of role-conditioned design are explored, including near-determinism not attributable to the stimulus strength, sensitivity to order of options in small gaps of evidence, and dependency of roles on presentation. The findings point to the conclusion that cross-lingual directional alignment alone does not constitute evidence for shared causation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.