acceptodds
Under review as a conference paper at ICLR 2027

InspectionBench : Evidence-Grounded Evaluation of Agentic Industrial Inspection

Abstract

With the rapid development of industrial intelligence and robot-assisted inspection, reliable physical-to-operational coupling from evidence acquisition and physical verification to downstream operational decisions has become a key bottleneck for deploying autonomous inspection and predictive-maintenance agents in real-world industrial environments. Existing benchmarks remain limited to static or single-modality observations, hand-specified scenarios, or outcome-oriented evaluation that fails to determine whether the required evidence was actually acquired. We introduce **InspectionBench**, an execution-grounded benchmark for agentic physical AI in industrial inspection, comprising 4,120 canonical episodes generated from 1,552 grounded worlds and 21 scenario templates across six industrial asset types, in which agents acquire visual, operational, and telemetry evidence with explicit provenance and execute inspections before making downstream operational decisions. A world-to-scenario compiler automates matched-episode generation through parameterized control of evidence visibility, modality and tool access, temporal validity, and guidance while preserving the underlying world and gold decision, adjudicated by domain experts where needed. We evaluate five complementary capabilities across five frontier models, finding that correct-action rates exceed evidence-supported success in 20 percentage points and exact physical-constraint identification up to 37.2 points, highlighting capability gaps that final task success alone can obscure. Code and evaluation artifacts are available in https://anonymous.4open.science/r/InspectionBench-3B71/README.mdCode.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.