Same Score, Different Blind Spots: What Object-Hallucination Benchmarks Can and Cannot Certify
Abstract
Object-hallucination benchmarks aggregate heterogeneous queries into pooled scores that do not reveal how false-positive risk varies across objects. We present an object-conditioned audit that assigns violation, unresolved, or certified verdicts and assesses whether benchmark allocation supports simultaneous object-wise certification. We evaluate eight open VLM checkpoints on POPE and DASH-B at a nominal 10% pooled false-positive budget. Kimi-VL and Molmo 2-O differ by 0.004 in pooled AUROC yet make 20 versus 49 false assertions on the same 147 POPE-popular cup negatives. The pair-level difference survives a selection-aware permutation test and persists under the models' native decisions. Using split-half held-out calibration and simultaneous exact bounds, we identify violations in 20 of 29 model benchmark cells in each direction, with 15 of 29 violating in both. A conservative calibration sensitivity analysis retains violations in a majority of cells in both directions. In the descriptive full-data census, no cell earns a simultaneous certificate covering all eligible objects. Under this convention, sparse allocation can preclude certification even with zero observed errors. An Open Images positive control on a checkpoint subset yields class-level certificates under frozen labels. A pre-specified readout gate voids Gemma 4's POPE cells, showing that statistical resolution alone does not establish valid model attribution. We recommend object-wise counts and uncertainty bounds, explicit unresolved outcomes, and validity checks alongside pooled scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.