acceptodds
Under review as a conference paper at ICLR 2027

VocabAudit: Auditing Inference-Vocabulary Sensitivity in Training-Free Open-Vocabulary Segmentation

Abstract

Open-vocabulary semantic segmentation (OVSS) is evaluated with a class-name list that directly controls test-time predictions, yet this list is often treated as fixed metadata. We study vocabulary sensitivity by freezing the images, model weights, prompt templates, and dense-inference pipeline while varying only the inference vocabulary. On a 300-image VOC-21 matrix, replacing an engineered vocabulary with plain dataset names reduces mIoU by 11.8–20.6 points. Filtered synonym substitutions cost up to 6 points and produce a 3.7-point spread across vocabulary draws. We then run VocabAudit V1.1 on 1,349 VOC images, three methods, six vocabulary-expansion conditions, and ten independent vocabulary draws per condition. Adding classes absent from the VOC-21 label set increases target-class mIoU by 4.49–12.75 points, with all 18 paired image-bootstrap intervals excluding zero, and reverses the cell-mean SCLIP–NACLIP ordering in every condition. At the same time, the expanded vocabulary substantially changes the predictions: with 200 added classes, pixel accuracy falls to 16.84–24.05%, and distractor labels receive 71.62–91.48% of benchmark-background pixels and 18.81–35.42% of foreground pixels. All-class mIoU also falls to 4.04–4.59, partly because the metric assigns zero IoU to label-set-absent classes. These results show that target class mIoU alone can obscure changes in the prediction interface and in pixel-level errors under vocabulary expansion. Reliable OVSS evaluation should therefore release the exact inference vocabulary, report sensitivity across controlled vocabulary draws, distinguish target-class from all-class accounting, and include an explicit pixel-flow ledger.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.