Error Detectability in Model Merging: Diagnosis and Constrained Repair
Abstract
Model merging can preserve a classifier's decisions while degrading the confidence ranking used to detect its errors. We isolate this effect on expert–merge agreement subsets, where predictions, and hence correctness labels, coincide: on the five CLIP ViT-B/32 tasks with at least ten shared errors, expert macro AUROC on these subsets is 86.4%, versus 69.2% for TIES. A three-class linear construction shows that parameter averaging can reverse a perfect error ranking while preserving every prediction, through an example-dependent Jensen gap in the competing logits. We then introduce PED (Preserving Error Detectability), a local task-vector correction that optimizes error ranking under margin constraints on protected correct fitting examples and retains only candidates whose macro fixed-label fitting AUROC is at least that of the base. With 128 fitting and 128 selection labels per task, PED raises eight-task held-out AUROC from 84.7% to 88.2% and accuracy from 71.7% to 73.4%, exceeding LMN and ProbSurgery by 2.8 and 2.3 points in direct paired comparisons. On ViT-L/14, it improves AUROC by 3.9 points on TIES and by 4.2 points on task arithmetic. With correctness labels held fixed, PED recovers 72.1% of the expert–TIES ranking gap on shared ViT-B/32 decisions, so its gains include better ranking of the base model's own errors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.