DFA-VLA: Dynamic Fine-Grained Alignment for Vision-Language-Action Models in Robotic Manipulation
Abstract
With the rapid advancement of robotic hardware and software technologies, Vision-Language-Action (VLA) models have become pivotal for mapping mul- timodal inputs to continuous robot actions. However, mainstream VLA models often suffer from poor modeling of fine-grained visual elements (e.g., occluded regions and small objects) and an over-reliance on static cross-modal attention, restricting adaptability in complex, open environments. A plausible but underex- amined failure mode is attention dispersion under visual clutter: when distractor objects are present, a dense visual prefix may spread cross-modal attention across irrelevant tokens and degrade manipulation precision even when the target object is clearly visible. To address these limitations, we propose DFA-VLA (Dynamic Fine-Grained Alignment for Vision-Language-Action Models), integrating two key components: (1) the Multi-scale Visual-Semantic Modeling (MVSM) mod- ule, which preserves global scene context while extracting segment-guided local object features; and (2) the Dynamic Fine-grained Alignment and Fusion (DFAF) module, which dynamically isolates task-critical object tokens via text-guided to- ken retrieval and learnable gating before cross-modal fusion into a frozen language backbone. In real-world evaluations on 22 manipulation tasks (220 trials, scored 1, 0.5, or 0 for success, partial success, or failure), DFA-VLA trained for 10,000 steps (5,000 on LIBERO followed by 5,000 on BridgeData V2) obtains a mean score of 78.6% (95% Wilson interval [72.8%, 83.5%]) against 71.1% for the pub- lished OpenVLA policy, scoring higher on 16 of the 22 tasks, tying on 5, and lower on 1 (paired sign test p < 10 ā3). In a paired 1,000-step clutter study on LIBERO- Spatial (32 scenes with 0/3/5/8 distractors), however, five-token prefixes, with or without segmentation, do not fit the demonstrations and reach only 3/32 and 1/32, while a LoRA-adapted dense control reaches 26/32 versus 22/32 for the published OpenVLA (exact McNemar p = 0.29). The two evaluations differ in budget, data, and scoring, and the physical evaluation has no matched dense or sparse control, so they are not directly comparable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.