SMA-VLA: Proprioception-Conditioned Semantic-Motion Alignment for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) to map visual observations and language instructions to robot actions, providing a promising paradigm for language-conditioned manipulation. However, strong in-distribution performance does not necessarily translate into robust behavior under distribution shifts, particularly when the robot starts from a different configuration, when an instruction is rephrased while preserving the same task semantics, or when visual layouts and distractors change. A key reason is that existing VLA models predominantly learn direct mappings from multimodal observations to expert actions, while how task semantics and current execution conditions influence the required motion remains insufficiently modeled. To address this problem, we propose SMA-VLA, a Vision-Language-Action framework centered on Proprioception-Conditioned Semantic–Motion Alignment (PC-SMA). The PC-SMA module incorporates the robot's current proprioceptive state into task conditioning and aligns the resulting task semantics with the coarse translational direction implied by expert demonstrations, thereby strengthening the otherwise implicit correspondence between task semantics and state-dependent motion. Multi-Level Context Enhancement integrates complementary information across VLM depths, providing richer semantic and spatial context for this alignment and reducing reliance on isolated layer representations. Patch-Focused Bridge uses the evolving action representation to emphasize action-relevant visual evidence, reducing interference from irrelevant scene factors and supporting more reliable motion prediction under visual shifts. Experiments on LIBERO, together with zero-shot transfer evaluations on LIBERO-Plus and LIBERO-Pro, show that SMA-VLA improves manipulation performance and enhances generalization under representative robot-state, instruction, and visual shifts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.