acceptodds
Under review as a conference paper at ICLR 2027

Spotting Is Not Inspecting: InspectBench for Coverage and Restraint

Abstract

Spotting one defect may flag an image, yet true inspection requires pinpointing them all. We tackle image-only defect inspection with vision–language models (VLMs), challenging a critic model to bound all visible defects or explicitly return an empty set for clean images. We present InspectBench, a cross-source benchmark designed to evaluate how models navigate the tension between defect coverage and false-positive restraint. The dataset spans 1,039 cross-reviewed images (321 clean controls) with 1,983 annotated regions across textual and visual defects, evaluated via a deterministic scoring protocol that penalizes duplicate hits and false alarms. Alongside the benchmark, we develop a reference inspection framework that combines content inventory, multi-scale evidence, and candidate verification, together with fine-tuned critic baselines spanning different supervision targets. Evaluating 17 model–procedure configurations across direct prediction and multi-stage inspection reveals a sobering baseline: region recall peaks below 34% at a lenient 0.1 localization threshold, leaving even the top-performing setup unable to touch 46.8% of defect regions. Furthermore, while selective regional review cuts false positives for some models, it backfires on others. These results expose a critical gap between high-level flaw detection and granular visual inspection, proving that evaluating coverage without false-alarm restraint gives a misleading sense of progress. Code, data, and benchmark tools will be released subject to licensing permissions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.