Which Layers Should Share? Functional Substitutability for Language Model Compression
Abstract
Cross-layer parameter sharing compresses a pretrained language model without shortening its computation graph, but it creates an assignment problem: which position-specialized components should share a physical base, and which member should initialize it? We introduce SUBSHARE, which treats this decision as directed functional role transfer. For each compatible donor-target pair, SUBSHARE measures the next-token loss change caused by substitution in the intact model, converts the resulting asymmetric matrix into medoid group costs, and reuses it across budgets with Greedy or mixed-integer planning. Our local-to-global analysis explains how these costs can select good recovered plans even when interactions affect their absolute loss. Across two Qwen3.5 scales and two Gemma models, SUBSHARE outperforms every topology-, spacing-, and weight-based assignment control in all 16 model-budget settings, beats SharVeT assignments in all 16, and beats Geo-Sharing assignments in 15. The gains extend to zero-shot tasks and persist under Basis Sharing parameterization. Functional substitutability thus provides a reusable decision principle for allocating a compact set of shared bases across a pretrained network.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.