Training-Free Protection of LLMs against Unauthorized Model Merging
Abstract
Model merging allows the specialized capabilities of a released language model to be incorporated into other models without authorization. We propose a training-free defense that adjusts model weights using parameter differences from compatible models with weaker target-domain performance. The guiding idea is to move the specialized model's parameters toward the boundary of a high-performing region in its loss landscape, retaining useful standalone performance while reducing the specialized capability preserved after subsequent merging. Global processing combines auxiliary-model weight differences to define the overall displacement. Local refinement applies alternating task-vector edits while preserving task-critical weights selected using calibration data. The procedure requires no additional gradient-based optimization or architectural changes. Experiments on mathematical and medical tasks show that the processed models retain most target-domain performance when used on their own, while contributing less capability through the evaluated merging algorithms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.