acceptodds
Under review as a conference paper at ICLR 2027

Calibrated Monitoring of Instruction–Execution Mismatch in Vision-Language-Action Policies

Abstract

Vision–language–action (VLA) policies can execute familiar motions while failing to follow the current instruction. Such instruction–execution mismatch can provide weak signals for monitors based on action uncertainty or novelty. To address this, we introduce **CALM**, a **CAL**ibrated **M**onitor that decodes instruction-conditioned spatial support from frozen VLA representations. Specifically, CALM constructs a Bayesian last layer (BLL) model to predict a distribution over end-effector locations together with an abstention output. After scalable variational training on successful execution endpoints and absent-target examples, we accumulate spatial support deficits from dispersed or abstaining predictions based on the trained BLL model at runtime and request assistance when the cumulative score exceeds a calibrated threshold. To obtain this threshold, we rely on split conformal prediction (CP), which entails held-out calibration episodes. Under episode-level exchangeability, split CP provides a finite-sample bound on the marginal probability of interrupting an otherwise successful episode. On LIBERO-10 with , , and OpenVLA, CALM achieves the highest detection rate, balanced accuracy, and time-weighted accuracy among evaluated methods within the empirical false-positive-rate budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.