UAM: A Dual-Stream Perspective on Forgetting in VLA Training
Abstract
Vision–language–action (VLA) models typically adapt a pretrained vision–language model (VLM) to robot control to transfer broad multimodal knowledge. Yet action fine-tuning can erode the competence that motivates this initialization, a loss we call the embodiment tax. Can a VLA reduce this loss while allowing its semantic backbone to adapt, without auxiliary vision–language co-training? Inspired by the recognition and visuomotor pathways of biological vision, we investigate a representational-bottleneck hypothesis: a semantic pathway that must also supply control-specific visual features may face competing adaptation demands. We study this hypothesis through the Unified Action Model (UAM), which couples a semantic expert, a parallel visual Dorsal Expert, and an action expert using mixture-of-transformers. Starting from aligned understanding–generation pretraining, we compare Dorsal initialization, visual versus query inputs, and visual-dynamics supervision. The strongest tested design pairs generative initialization with future-observation prediction from robot trajectories, providing a complementary learning target while all three experts remain trainable. Under matched initialization and action-training settings, UAM retains 96.7% of pretrained scores across four reported multimodal metrics, versus 54.8% without the Dorsal pathway, while also improving OOD action performance by 35.0% with no auxiliary vision–language replay, VLM freezing, or gradient blocking is used during adaptation. After training on only 3,000 real-robot trajectories, UAM achieves the highest average success among the compared variants on semantic generalization tasks involving unseen objects, novel target compositions, and instruction variation. Together, these results support aligned generative pretraining and visual-dynamics supervision as a means to improve the retention–control trade-off in unified multimodal models. Demos can be found at https://sites.google.com/view/unified-action-modelour anonymous website.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.