acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language-Action Models Lose Information as They Fail

Abstract

Vision-language-action (VLA) models are increasingly deployed on physical robots, where failures can cause irreversible harm. This calls for failure detection that generalizes across policies and explains why a rollout fails. We start from an information-theoretic premise: effective information transfer is a prerequisite for achieving a goal. Formalizing VLA control as a closed-loop information pipeline, we derive eight information-theoretic metrics over states and actions and select three complementary ones, the Triple Information-theoretic (Tri-Info) metrics: action entropy, temporal action coherence, and action–state coupling. They characterize a healthy policy as one that acts with moderate entropy grounded in both its history and its observations, while failure violates at least one of these conditions. A GRU-based detector predicts failure from how these signals evolve. Across various VLA models, simulated benchmarks, and real-world tasks, Tri-Info matches the strongest baselines in domain and transfers across architectures, environments, and the sim-to-real gap without target-domain failure labels, reaching over 70% accuracy on real-world tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.