Deciding Which Detections to Release: BlockCalibration for Open-Vocabulary Detection
Abstract
Open-vocabulary detectors let users locate objects through text, but a confident prediction for an absent category can become a false alert or an incorrect training label. Reliable use therefore requires deciding which detections to release while retaining useful objects. This decision is complicated by the detector itself: each query produces a dependent set of candidates whose size changes with the image, query, and score filtering. Pooling these candidates for calibration gives queries with more proposals greater weight and can fail to control the average error across queries. We propose block calibration, a postprocessing framework that aligns calibration with this query-level objective. The method keeps each image and its queries together, averages false-detection evidence within queries and across the image, and uses the resulting reference with the e-value Benjamini–Hochberg rule to select detections. Under exchangeability of complete images and their queries, it controls the expected false fraction averaged across queries, allowing score-dependent candidate counts and arbitrary dependence within an image. A maximum-reference variant also controls the expected false fraction in an image’s combined output. We further characterize the exact release condition, explaining how evidence strength, candidate count, and rank determine which objects survive. Experiments with GroundingDINO and YOLO-World on COCO, LVIS, and Open Images examine both high-precision release and its cost in object recovery. On COCO/GroundingDINO-tiny, the learned-evidence rule recovers 16.51% of objects at 98.12% pooled precision, more than doubling the recovery of binary evidence at the same nominal error level. Filtering and prompt-bank studies further show how candidate generation shapes the usefulness of error-controlled release.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.