Do LoRA Merge Diagnostics Score the Deployed Model?
Abstract
A LoRA merge diagnostic can fail at three distinct stages: the update it scores, the curvature it estimates, and the endpoint it is intended to predict. We analyze these mismatches while holding candidate construction and evaluation roles explicit. An exact factor-covariance identity separates factor aggregation from scalar-equivalent correction and rank truncation. For sequence-balanced Fisher estimation, probe-count correction removes aggregation bias but retains cross-probe variance; categorical marginalization removes label-sampling variance. A five-adapter Pythia-1.4B case study exhibits a deployed-to-reference norm ratio and task-dependent local-quadratic miscalibration. Controlled studies use twenty independently trained banks per panel, two architectures, and a fixed twenty-four-candidate library. Replacing uniform logit drift with specialist-probability weighting lowers mean worst-task KL regret from to in a nonlinear classifier and from to in causal attention. This preservation gain does not establish a native-accuracy gain: weighted-minus-uniform attention accuracy is percentage points, with paired interval . Selection for native retention using selection labels improves accuracy over KL selection by percentage points . The results support matching diagnostics to deployed updates and preservation geometry, while evaluating native utility separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.