acceptodds
Under review as a conference paper at ICLR 2027

Where Is Failure Predictable? Stage-Wise Attribution in Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models map observations and instructions to robot actions, but they still fail in ways that can damage objects or leave the robot in unrecoverable states. A VLA processes each step in three phases: it aligns the scene with the instruction, reasons over the aligned representation, and decodes an action. A failure can originate in any of these phases, yet existing failure detectors return a single score for the whole policy. We derive a theoretical characterization of uncertainty and failure predictability across the VLA processing pipeline. In particular, we establish an exact decomposition of predictive action entropy into alignment, reasoning, and decoding terms with nonnegative population values, and prove that failure predictability likewise decomposes into nonnegative gains from phase-restricted scores. Motivated by our theoretical results, we develop STAGE, a computational framework for phase-wise failure prediction. STAGE extracts alignment and reasoning signals from their stage representations, computes decoding uncertainty directly from the decoder, and integrates the three signals to predict failure in a single VLA forward pass. Extensive experiments demonstrate the effectiveness of STAGE across three VLA families (OpenVLA, , and -FAST), four LIBERO suites, and SimplerEnv. STAGE achieves the highest AUROC in most settings, outperforming the strongest baseline by up to points, and outperforms all evaluated baselines on real-robot Agilex rollouts. Stage-targeted perturbations further validate the phase-wise attribution: each score responds to its corresponding stage while remaining unaffected by perturbations to later stages, revealing where failure signals arise in the VLA pipeline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.