The Geometry of Weight-Tying Sensitivity in Transformers
Abstract
Transpose tying changes both a learned operator and the residual trajectory on which later computation acts. We develop a finite geometric account of these two effects in pre-normalized Transformers. Positive channel scaling yields a unique balanced SwiGLU representation and a weighted-adjoint tying operation that is invariant under function-preserving channel rescaling. Exact gated derivatives identify a rank-two contribution to normalization-induced rotation, while radial–directional coordinates distinguish residual magnitude from the angular change produced by an update. For a fixed pre-RMSNorm suffix at input scale , we derive explicit coefficients for its logit and output-KL response, together with finite-radius bounds and direction-resolved endpoint certificates. Public-model interventions show large propagation differences after local edit-magnitude matching. A complete two-model, six-text strength panel tests the balanced edit itself: its coordinate consistency coexists with substantial functional changes, including a low-amplification regime. A separate two-model reference-precision experiment computes the suffix coefficients without fitting output responses: at , the six cut-level forward-KL ratios to prediction range from to . The theory links a well-defined editing operation to residual geometry and quantitatively testable suffix effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.