Mixture-of-Checkpoints: Improving Initial-State Coverage of Vision-Language-Action Policies via Online Routing
Abstract
Post-training for vision-language-action policies usually keeps the checkpoint with the highest aggregate success rate and discards the rest. This selection ignores whether different checkpoints cover different initial states. We study whether checkpoints that do not improve aggregate success can still improve initial-state coverage. We post-train OpenVLA-7B with Direct Preference Optimization on action preference pairs derived from Monte Carlo tree search value estimates; each post-training run yields one checkpoint. The resulting checkpoints match or underperform the base model on LIBERO. However, a two-round per-state audit shows that they succeed on different initial states. Together, they cover most of the initial states where the base model fails. Therefore, we deploy the base model and all checkpoints as a policy portfolio, which we call Mixture-of-Checkpoints, and add a router that selects one policy per episode. The router is a per-state contextual bandit: the initial-state index is the context, episode success is the reward, and an exponential moving average updates a per-state score table. Across eight LIBERO tasks, routing matches or outperforms the base model. On stable tasks with headroom, it adds 8 to 19 successful episodes per 100 episodes. Two tasks selected after the method was fixed confirm predictions recorded in advance. Ablations show that the gain requires online per-state measurement, and a two-round audit predicts which tasks benefit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.