MMA: A Simple and Flexible Multi-Model Framework for Action
Abstract
Vision-language models (VLMs) and vision-language-action models (VLAs) provide diverse pretrained representations for robot action learning, shaped by different data, architectures, and objectives. Combining their strengths is challenging because their interfaces differ, while running multiple models increases inference cost. We introduce MMA, a framework that jointly learns a fused policy and single-source branches that can be deployed independently. Learnable action queries and branch-local adaptation connect each source to a shared action interface and action head. An Action-Centric Mixture-of-Transformers (MoT) fuses branch representations, and a stop-gradient fused-query teacher transfers information to the branches through query alignment and relational matching. The resulting policies support either full-fusion or single-branch inference. On LIBERO, RoboTwin 2.0, and real-world manipulation tasks, the fused policy achieves average success rates of 99.2%, 93.1%, and 93.8%, respectively, with absolute success-rate gains of 1.7%, 5.5%, and 15.0% over the strongest matched single-source baselines. All three MMA branches also improve over their matched baselines across these evaluations. Each branch runs only its own source and the shared action head, providing a lower-cost deployment choice alongside fusion. MMA reuses pretrained components without access to their original pretraining data or additional large-scale pretraining. MMA offers a reference point for future work on combining pretrained VLMs and VLAs while balancing policy performance and deployment cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.