acceptodds
Under review as a conference paper at ICLR 2027

VLA-Hallucination: Diagnosing and Mitigating Hallucination Failures in Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models have recently become an important technique for general-purpose robotic manipulation. However, existing VLA models still exhibit clear limitations in generalization to unseen tasks. We find that VLA task failures can be categorized into manipulation failures and hallucination failures. The latter refers to cases where the model-generated actions are inconsistent with the target task or the surrounding environment, and constitutes a major factor limiting VLA generalization. Further analysis reveals that an important cause of hallucination failures is that task-relevant visual-semantic information preserved in the upstream VLM is not effectively utilized during downstream action generation, causing the generated actions to deviate from the current task semantics. To address this issue, we propose Action-to-Visual Semantic Alignment (AVSA). AVSA introduces a lightweight visual-semantic reconstructor that reconstructs visual-semantic information conditioned on the action, online measures the alignment between Action Expert outputs and VLM visual semantics, and uses this signal to correct actions during Flow Matching, thereby aligning action generation with the upstream visual-semantic understanding. AVSA can be integrated with multiple mainstream VLA models without retraining the VLA backbone, and consistently reduces hallucination rates while improving task success rates across diverse unseen settings, including environment changes, semantic changes, and paraphrased instructions. For example, on LIBERO-Para, AVSA improves the task success rate of by 7.1%. Overall, our results demonstrate that explicitly diagnosing and aligning action generation with visual semantics provides an effective approach to mitigating hallucination failures and improving the generalizability of VLA models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.