Introspective Signals for Failure Prediction in Vision-Language-Action Models
Abstract
Existing approaches for failure prediction in VLAs rely on information derived from policy inputs, outputs, and latent representations. However, the predictive value of relationships among internal signals (such as attention, value vectors, and hidden states) and their evolution during execution remains underexplored. We hypothesize that these patterns can support early and accurate prediction of a VLA’s eventual failure. We construct hypothesis-driven introspective signals organized around four categories – world sensitivity, temporal consistency, attention–attribution agreement, and internal coordination – and learn temporal patterns over them that align with with episode-level outcome labels. Across three policy families and two LIBERO-based task settings, our models predict failures earlier and more accurately than existing methods and generalizes to unseen tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.