When the Whole Exceeds the Parts: Constructing Multi-Ability LLMs via Component-wise Weighted Fusion
Abstract
Specialized language models often perform poorly on tasks that require capabilities from different domains. We introduce Component-wise Weighted Fusion (CWF), a model-merging method that learns how much each expert contributes to individual feed-forward network (FFN) neurons and attention heads. An FFN coefficient is shared across the neuron’s Gate, Up, and Down projections. Attention coefficients are shared by corresponding query and output slices, with group-averaged coefficients for shared key/value heads. The expert weights remain frozen while the coefficients are trained on mixed-domain data; the resulting weights form one dense model with no additional inference-time experts. We evaluate mathematics in Indonesian and Thai, coding with mathematics, and coding with multilingual knowledge using Llama and Qwen models at the 8B scale. Without target-language mathematics examples in mask training, CWF improves Indonesian and Thai MATH accuracy over the stronger input expert by 13.98 and 15.70 percentage points. On Qwen Code+Math, it matches the best reported merging baseline on each coding benchmark while retaining strong mathematical performance. We jointly evaluate target-language gains and source-domain retention. Component and learning-rate ablations favor joint FFN/attention fusion with separate learning rates, while mask and activation analyses expose heterogeneous expert contributions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.