acceptodds
Under review as a conference paper at ICLR 2027

CASH: Calibrated Self-Healing Vision-Language-Action Policies for Embodied Agents

Abstract

Action-chunking vision-language-action (VLA) policies predict multiple future actions and execute them without incorporating new observations, local execution errors can thus accumulate over the open-loop horizon. To address this issue, we present CASH (Calibrated Self-Healing), a lightweight failure-detection and self-healing framework that keeps the VLA frozen and consists of detection and self-healing. The detection module contains an execution anomaly detector (EAD) and a policy anomaly detector (PAD). EAD tests whether an executed action is consistent with the observed latent transition. Its score combines forward prediction error, directional disagreement, and inverse-dynamics reconstruction error; a conformalized quantile threshold with a one-sided CUSUM (cumulative sum) that converts the score into a state-conditioned trigger for persistent deviations. PAD tests whether a proposed action is plausible under the current observation and instruction, using successful actions and several families of counterfactual actions to train a value model. When EAD triggers, CASH computes the latent displacement that should have occurred but has not yet occurred, and adjusts the action chunk through a lightweight dynamics model. When PAD triggers, CASH shortens the action chunk and triggers earlier re-planning; repeated low scores activate a window reset for severe out-of-distribution (OOD) conditions. The detection modules are trained on cached latent features using successful trajectories only, without failure data or manual failure labels. The self-healing module does not require additional training. CASH can be integrated into different VLA models without retraining the VLA backbone, while preserving the efficiency of action chunking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.