Efficient Untying of Weight-Tied Language Models via Asymmetric Geometric Decoupling
Abstract
Weight tying reduces vocabulary-layer parameters, but forces input embeddings and output unembeddings to share one vocabulary geometry. Modern language models increasingly untie these matrices at larger scales, raising the question of whether parameter-constrained models can relax tying without paying the full cost of an independent vocabulary matrix. We diagnose this trade-off through training dynamics and embedding geometry across public model families. Tied shared matrices are consistently closer to untied output unembeddings than to input embeddings; output spaces also show stronger cross-model cosine alignment and stronger mean/rank-1 spectral structure. These patterns match dense softmax gradients on the output side versus sparse lookup gradients on the input side, suggesting that hard tying makes the shared matrix output-dominated and constrains independent input-side geometry. We propose Asymmetric Geometric Decoupling (AGD), which keeps a shared vocabulary matrix on the output side while adding small structured freedom to the input side through bias correction and hidden-dimensional low-rank deformation. Input-side AGD improves tied baselines in MiniMind and FineWeb-Edu pretraining, including a 0.6B configuration trained from scratch on 20B tokens, with minimal parameter overhead. Its gains persist in held-out assistant-token likelihood after full-parameter SmolTalk fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.