acceptodds
Under review as a conference paper at ICLR 2027

From Rollouts to Failure Modes: Behavioral Diagnosis of Vision-Language-Action Models

Abstract

Task success is the standard measure for evaluating Vision-Language-Action (VLA) models, but it provides limited insight into the failures underlying aggregate performance. We introduce embodied failure attribution, a framework that extends success-rate evaluation with behavioral diagnosis. Using Bayesian inference over stage-specific physical interaction evidence from closed-loop rollouts, the framework assigns posterior probabilities to Planning, Grounding, and Action (PGA) failure modes without requiring access to model internals. This provides a common diagnostic representation for characterizing failure structure, tracking its evolution, and informing targeted interventions. Across diverse VLA architectures and task settings, we find that Action is a recurring source of failure, while the relative contributions of PGA failure modes vary across models and tasks. During training, the framework distinguishes improvements in overall performance from persistent failure modes: Qwen3VL-PI achieves higher task completion on CALVIN, yet Action remains the dominant failure dimension at every evaluated checkpoint, while Grounding attribution declines overall and Planning attribution fluctuates. Under unseen-object conditions, attribution identifies Action as the dominant failure mode, providing additional insight into the observed generalization failures. Finally, an intervention study illustrates how diagnosis can inform model improvement. Action-level correction raises success on an Action-limited task from 51% to 61%, but lowers success on a task with a different failure profile from 65% to 55%, suggesting that the benefit of a correction depends on the failure it addresses. Together, these results establish behavioral failure attribution as a diagnostic complement to aggregate performance evaluation, supporting the analysis of VLA systems across model comparison, training, and generalization, while providing a basis for targeted improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.