acceptodds
Under review as a conference paper at ICLR 2027

Introspective Signals for Failure Prediction in Vision-Language-Action Models

Abstract

Existing approaches for failure prediction in VLAs rely on information derived from policy inputs, outputs, and latent representations. However, the predictive value of relationships among internal signals (such as attention, value vectors, and hidden states) and their evolution during execution remains underexplored. We hypothesize that these patterns can support early and accurate prediction of a VLA’s eventual failure. We construct hypothesis-driven introspective signals organized around four categories – world sensitivity, temporal consistency, attention–attribution agreement, and internal coordination – and learn temporal patterns over them that align with with episode-level outcome labels. Across three policy families and two LIBERO-based task settings, our models predict failures earlier and more accurately than existing methods and generalizes to unseen tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.