SpurDiag-Bench: Diagnosing Spurious Correlations in Vision-Language-Action Models
Abstract
Despite strong benchmark performance, there is growing consensus that many Vision-Language-Action (VLA) models still fail in real-world scenarios, suggesting that benchmark success overstates their real-world generalization. While most previous work attributes it to a data-scale problem, we observe that it can also stem from a spurious correlation phenomenon: through shortcut learning, models associate targets with incidental cues rather than desirable semantics and task logic. To systematically diagnose this phenomenon, we introduce SpurDiag-Bench. SpurDiag-Bench distinguishes itself from existing VLA benchmarks through diagnostic depth: it revisits known classic VLA tasks at fine granularity and reveals that VLA failures do not require harder/broader tasks - simple changes in task configuration alone are sufficient to expose failures hidden behind strong benchmark performance. Precisely, we introduce a chain analysis that decomposes the model along its perception-reasoning-action chain and applies causal interventions to separately examine the grounding of visual-language and actions. Across experiments on 12 state-of-the-art open-source VLAs, we report that spurious correlations are pervasive (even substantial for the best-performing model, ). We further fine-tune on representative models, investigating the mitigation of spurious correlations: some models become less reliant, while others retain persistent shortcut reliance. These findings caution against assuming that shortcut reliance can be readily corrected through additional training. Key properties are also identified in robust VLA models featured for countering spurious correlations (e.g., action-expert architecture, diverse pretraining), providing guidance for future model design for the community. The benchmark and source code will be released, and the leaderboard is available at https://vla-page.netlify.app.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.