acceptodds
Under review as a conference paper at ICLR 2027

SpurDiag-Bench: Diagnosing Spurious Correlations in Vision-Language-Action Models

Abstract

Despite strong benchmark performance, there is growing consensus that many Vision-Language-Action (VLA) models still fail in real-world scenarios, suggesting that benchmark success overstates their real-world generalization. While most previous work attributes it to a data-scale problem, we observe that it can also stem from a spurious correlation phenomenon: through shortcut learning, models associate targets with incidental cues rather than desirable semantics and task logic. To systematically diagnose this phenomenon, we introduce SpurDiag-Bench. SpurDiag-Bench distinguishes itself from existing VLA benchmarks through diagnostic depth: it revisits known classic VLA tasks at fine granularity and reveals that VLA failures do not require harder/broader tasks - simple changes in task configuration alone are sufficient to expose failures hidden behind strong benchmark performance. Precisely, we introduce a chain analysis that decomposes the model along its perception-reasoning-action chain and applies causal interventions to separately examine the grounding of visual-language and actions. Across experiments on 12 state-of-the-art open-source VLAs, we report that spurious correlations are pervasive (even substantial for the best-performing model, ). We further fine-tune on representative models, investigating the mitigation of spurious correlations: some models become less reliant, while others retain persistent shortcut reliance. These findings caution against assuming that shortcut reliance can be readily corrected through additional training. Key properties are also identified in robust VLA models featured for countering spurious correlations (e.g., action-expert architecture, diverse pretraining), providing guidance for future model design for the community. The benchmark and source code will be released, and the leaderboard is available at https://vla-page.netlify.app.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.