acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language-Action Model Already Has Attention Heads for Path Deviation Detection

Abstract

Vision-Language-Action (VLA) models have demonstrated strong potential in navigation tasks by predicting semantic actions and reasoning over complex linguistic and visual contexts. However, they are fundamentally hindered by visual-reasoning hallucinations causing trajectory deviations. Addressing this issue has conventionally required training external critic modules or relying on complex uncertainty heuristics. In this work, we discover that monitoring specific attention heads within a frozen VLA model detects path deviations without incurring additional computational overhead. We term these Navigation Heads, as they inherently capture the spatiotemporal correlations between historical visual sequences and linguistic instructions. Using these heads, we propose an intuitive, training-free anomaly-detection framework that monitors their signals to detect hallucinations in real time. Surprisingly, among over a thousand attention heads, a combination of just three is sufficient to achieve a 65.1 % anomaly detection rate (recall) and a 76.4 % F1 score in unseen environments. Furthermore, we couple this high-level detection logic with a low-level Reinforcement Learning (RL) policy for local obstacle avoidance, ensuring stable navigation and executing a direct safe recovery upon anomaly detection. Ultimately, deploying this entire framework zero-shot onto our customized physical robot demonstrates its practical robustness. All source code will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.