DoOver: Detect Failure from Demos, Do It Over On Demand
Abstract
Current robot policies often continue acting obliviously after execution goes wrong. We introduce DoOver, which couples lightweight, policy-native failure detection with on-demand recovery mechanism. With a M-parameter head that learns jointly with the policy from timestamps in expert demonstrations and reuses policy's visual features, DoOver tracks changes in predicted progress to detect stalls and regressions. It requires no failure data, pretrained foundation models, separate detector training, or runtime action resampling. Across simulated tasks with flow-matching and diffusion backbones and detection baselines per backbone, DoOver leads the eligible baselines in aggregate average precision across task-calibrated detection, alarm-threshold sweeps, and threshold generalization, with lower failure-detection compute than the sampling-based detectors STAC and ACE. Beyond detection, we establish a systematic benchmark of failure-triggered recovery for frozen policies, comparing seven operators spanning five inference-time intervention families. Among them, failure-triggered feedback world model improves recovery rates over plain replanning by / pp and over normal execution by / pp for flow-matching/diffusion policies, while avoiding the high cost of guided sampling during normal execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.