A Controlled Benchmark for Weight-Space Model Merging Across Medical Imaging Domains
Abstract
Consolidating medical-imaging specialists into a shared model requires preserving performance across distinct imaging domains. Weight-space model merging offers a way to combine independently fine-tuned encoders without retraining, but comparisons between merging methods are sensitive to validation protocols. Official evaluation scripts for TSV-Merge and isotropic merging use wider scaling-coefficient ranges for selected spectral methods than for other tunable baselines. We introduce a controlled benchmark spanning dermoscopy, chest radiography, and histopathology, with three binary tasks and training images per specialist. We compare 11 methods across seven vision-transformer backbones and three seeds using a shared scaling-coefficient grid. The best observed merger reaches mean AUROC against for task-aware specialist routing, with tuned Task Arithmetic close behind at . We also introduce Gram-LS, a closed-form method based on task-vector projection constraints that performs comparably to tuned summation. Within-method validation tuning improves Gram-LS and Task Arithmetic by approximately – AUROC, exceeding the spread among five leading tuned methods, while search bounds can change method rankings. Iso-C and Iso-CTS fall – AUROC below tuned summation. These findings establish tuned summation as a strong baseline in this setting and show that specialist references and validation protocols are essential to assessing the performance cost of consolidation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.