Path-Constrained Model Merging for Robust Vision-Language Models
Abstract
Vision-language models (VLMs) deployed in real-world settings need to handle both natural distribution shifts and adversarial perturbations. Model merging provides an efficient way to combine these two capabilities without additional training. However, we find that existing merging methods, while achieving clear OOD gains, do not reliably preserve the robustness already present in an adversarially robust model. We further observe that merging affects perturbation sensitivity differently across modules, and that correcting only a small number of sensitive parts can recover a large portion of the lost robustness. In this paper, we propose Path-Constrained Model Merging (PACOMERGE), a training-free and data-free post-merging approach. PACOMERGE operates along the merge path produced by an existing model merging method and uses the perturbation sensitivity of the adversarially robust model to define a constrained region along this path. Updates that remain within the constraint are kept unchanged, while those that exceed it receive only the necessary correction. We further show that the points satisfying the constraint form a continuous interval along the merge path, allowing PACOMERGE to retain as much of the original merge as possible while making only the minimum correction needed to recover robustness. Across five base mergers, PACOMERGE raises average robust accuracy to 18.24–27.78% on ImageNet and four distribution shifts, with consistent gains on zero-shot and LLaVA tasks while preserving multi-task merging performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.