acceptodds
Under review as a conference paper at ICLR 2027

PreMerge for VLA: Building a Stronger Model by Merging Models with a Shared Base

Abstract

Vision-language-action (VLA) models commonly initialize from a single pretrained vision-language model (VLM), although models from different training stages may retain complementary knowledge. We propose PreMerge for VLA to combine VLMs with a shared base before target-task action training. First, we align their vision-language parameters and assign source weights from the magnitude and direction changes of the original weights relative to the base. A sign-consistency mask removes conflicting parameter updates. Second, we use a small set of target images and instructions, without action annotations, to calibrate the merged model through multilayer feature distillation. Source parameters remain frozen, and only tensor-level merging coefficients are optimized. Using the aligned VLM components of Qwen3-VL, Cosmos-Reason2, and GR00T-N1.7, PreMerge achieves average success rates of 98.9% on LIBERO and 92.7% on RoboTwin 2.0, exceeding the best single-source initialization by 1.0 and 3.2 percentage points, respectively. On four real-world manipulation tasks, it achieves 90.0%, compared with 82.5% for the strongest evaluated external baseline. The merged parameters are folded into one VLM, so subsequent action training and inference require no additional source branches. These results demonstrate the value of reusing models from different training stages to construct VLA initializations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.