Semiclassical Geometry of Attention Switching
Abstract
When one attention head is pushed toward a competing evidence span, do baseline quantities identify useful directions for controlling the answer margin? For a fixed-key, two-branch logit shift, the head's output change splits exactly into a branch-switching term and a total-evidence-mass term (a specialization of softmax differentiation, used as coordinates). Switching geometry identifies a task-enriched control subspace. Across eight Qwen3-8B examples (archived custom path), the three-dimensional switching normal subspace retains a mean of squared task-gradient norm versus for the median matched-random subspace (paired difference , 95% interval ; uses each example's gradient). Baseline-only prediction holds. On 50 ConflictQA examples and eight Gemma heads, finite attention recomputation from the baseline gives pooled versus for local linearization. In an exploratory native Qwen3-8B cohort (a 32-fact prefix of a 200-fact ConflictQA test set, path gate bypassed), a predictor frozen from baseline forward and backward passes achieves small median local-regime error () for all three steering arms—switching, fixed PCA basis, and static mean direction: geometrically grounded directions are collectively useful. A switching-specific advantage is not established. On those 32 facts, switching minus PCA at is and switching minus static is (not a test of equivalence); the 200-fact cohort favors switching over the fixed basis only. Switching geometry thus locates useful control directions but is not a privileged axis; path and attribution limits are in Section sec:discussion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.