acceptodds
Under review as a conference paper at ICLR 2027

The Geometry of Weight-Tying Sensitivity in Transformers

Abstract

Transpose tying changes both a learned operator and the residual trajectory on which later computation acts. We develop a finite geometric account of these two effects in pre-normalized Transformers. Positive channel scaling yields a unique balanced SwiGLU representation and a weighted-adjoint tying operation that is invariant under function-preserving channel rescaling. Exact gated derivatives identify a rank-two contribution to normalization-induced rotation, while radial–directional coordinates distinguish residual magnitude from the angular change produced by an update. For a fixed pre-RMSNorm suffix at input scale , we derive explicit coefficients for its logit and output-KL response, together with finite-radius bounds and direction-resolved endpoint certificates. Public-model interventions show large propagation differences after local edit-magnitude matching. A complete two-model, six-text strength panel tests the balanced edit itself: its coordinate consistency coexists with substantial functional changes, including a low-amplification regime. A separate two-model reference-precision experiment computes the suffix coefficients without fitting output responses: at , the six cut-level forward-KL ratios to prediction range from to . The theory links a well-defined editing operation to residual geometry and quantitatively testable suffix effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.