AttnCal: Native Attention Construction and Visual Calibration for Multimodal Model Merging
Abstract
Attention logits depend jointly on query and key (Q/K) projections, so preserving each projection separately need not preserve their interaction. In multimodal merging, this interaction must also process visual features from a consolidated interface. AttnCal addresses both requirements through native attention construction and visual-conditioned calibration. It reconstructs joint rotary-position operators under the original head-width constraint, targeting the worst expert-relative reconstruction error within each head. The construction uses expert weights alone, provides computable objective bounds, and controls attention-logit error for fixed hidden inputs uniformly over relative position. AttnCal calibrates the constructed Q/K factors under a fitted visual map, then fixes the language weights while adapting a shared visual residual, retaining one static model and the original visual interface. Across four reproduced benchmark–backbone protocols, AttnCal achieves the highest final average score (FAA) among merged baselines. On six UCIT tasks with LLaVA-1.5-7B and 3,000 examples per task, it reaches 76.07 FAA, 82.72 continual average score (CAA), and 0.33 backward transfer (BWT), compared with 72.15, 81.60, and for the strongest FAA baseline. These results connect a native-feasible attention target to the visual state under which the merged model is calibrated and deployed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.