acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Protection of LLMs against Unauthorized Model Merging

Abstract

Model merging allows the specialized capabilities of a released language model to be incorporated into other models without authorization. We propose a training-free defense that adjusts model weights using parameter differences from compatible models with weaker target-domain performance. The guiding idea is to move the specialized model's parameters toward the boundary of a high-performing region in its loss landscape, retaining useful standalone performance while reducing the specialized capability preserved after subsequent merging. Global processing combines auxiliary-model weight differences to define the overall displacement. Local refinement applies alternating task-vector edits while preserving task-critical weights selected using calibration data. The procedure requires no additional gradient-based optimization or architectural changes. Experiments on mathematical and medical tasks show that the processed models retain most target-domain performance when used on their own, while contributing less capability through the evaluated merging algorithms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.