SCouT: Softly Coupled Fine-Tuning for Preserving Specialization while Improving Mergeability
Abstract
The pre-training-fine-tuning paradigm is widely used to adapt large-scale models to downstream tasks, but maintaining separate task-specific models incurs substantial storage and deployment overhead. Model merging addresses this by combining multiple specialists into a single multitask model. However, independently optimized specialists can be poorly compatible, leading to a substantial drop in merged-model performance. We propose Softly Coupled Fine-Tuning (SCouT), which introduces merge-aware coupling during fine-tuning to improve mergeability while preserving task-specific performance. The coupling strength interpolates between independent fine-tuning and conventional multitask learning, recovering independent fine-tuning at zero coupling and approaching a shared multitask solution under strong coupling. Our analysis identifies a favorable intermediate regime in which weak coupling improves post-merge risk linearly in the coupling parameter while affecting specialist risk only quadratically, explaining why mergeability can improve much faster than specialization degrades. Experiments across model architectures and merging rules show that SCouT substantially improves post-merge performance while retaining strong specialists. These results suggest that weak cross-task coupling can steer fine-tuning toward more compatible solutions while still largely preserving their standalone specialization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.