acceptodds
Under review as a conference paper at ICLR 2027

FailBench: Benchmark for robotic manipulation failure detection

Abstract

Vision-language models (VLMs) are increasingly used to determine whether a robot manipulation attempt has succeeded. These judgments can serve as reinforcement-learning rewards, training-data filters, policy-ranking signals, or triggers for retry, making reliable failure detection critical for robot learning and evaluation. However, existing benchmarks provide limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts from 14 public sources, including 12 real-world and 2 simulated datasets, using the original outcome labels provided by their respective sources. Six of the real-world sources were originally collected for policy evaluation, reward modeling, or general data collection rather than failure detection, and 76% of failures occur naturally rather than being deliberately constructed. Evaluating 13 VLM-based detectors, we find that the best model achieves only 0.77 mean balanced accuracy, indicating substantial room for improvement. Models fine-tuned specifically for robot failure detection consistently underperform general-purpose VLMs, with the exception of the smallest general-purpose model, and also underperform their corresponding pretrained models. We further find that performance depends more strongly on the visual evidence required to determine the outcome than on the robot or task domain: detection approaches saturation when success is determined by observable object motion, but approaches chance when success depends on establishing contact, with no model exceeding 0.60 balanced accuracy on contact-intensive assembly tasks. Increasing reasoning effort does not alleviate this bias, as incorrect predictions tend to receive longer reasoning traces. Model-level interventions provide little improvement, while an input-level intervention is effective: spatially localizing the outcome-relevant region and cropping the input improves the strongest detector by 2.3 percentage points without additional training. We release FailBench and the accompanying evaluation harness to facilitate reproducible evaluation of robot failure detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.