acceptodds
Under review as a conference paper at ICLR 2027

MOSAIC: Multi-Policy Sampling and Collaborative Test-Time Adaptation for Vision-Language-Action Models

Abstract

Pretrained vision-language-action (VLA) policies integrate visual observations and natural-language instructions to generate executable robot actions, provid- ing a strong foundation for language-conditioned visuomotor control. However, large-scale pretraining and task-specific adaptation do not guarantee consistent deployment success. A single policy may produce different execution outcomes across stochastic action samples, while different policies may exhibit complemen- tary successes and failures across initial states within the same task set. These variations create opportunities to discover successful behaviors during deploy- ment. Yet a successful execution does not automatically become transferable to other policies or reliably reproducible by the policy that generated it. We there- fore investigate how collaborative adaptation can turn successful deployment-time experience into transferable and reproducible policy behaviors. We propose MOSAIC, a Multi-pOlicy Sampling And Inter-model Collaboration framework for test-time adaptation of vision-language-action models that learns from successful trajectories generated by multiple policies and sampling seeds us- ing environment-provided task-completion feedback. Starting from the same ini- tial simulator state, participating policies perform rollouts with different sampling seeds and exploit successful experience through two complementary mechanisms. Cross-policy action distillation aligns other policies’ action predictions with suc- cessful actions at the corresponding observations, facilitating behavioral transfer across policies. Within-Policy Sampling Consolidation incorporates a policy’s own successful actions into its native generative training objective, encouraging successful behaviors to recur across stochastic samples. Both mechanisms up- date only a designated module near each policy’s action output, requiring neither additional expert demonstrations nor additional trainable modules. We adapt the policies on a subset of initial states for each task, then freeze the resulting parameters and evaluate on disjoint held-out initial states. Experimental results show that across four manipulation benchmarks and 52 policy pairs, CAD improves 103 of 104 policy-side evaluations with an average gain of 3.03 per- centage points, while WTC independently improves average task success by 2.91 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.