HIVE: Holistic Image-Adaptive Vocabulary Evolution for Training-Free Dense Vision-Language Inference
Abstract
Benefiting from pretrained vision-language models, training-free open-vocabulary semantic segmentation (OVSS) predicts pixel-level labels for arbitrary categories without task-specific optimization. Existing training-free OVSS methods commonly perform dense prediction using independently processed local views and a fixed set of category names. However, this paradigm suffers from two limitations. First, processing local views independently prevents the model from capturing cross-view dependencies, limiting contextual reasoning for dense prediction. Second, fixed category names describe semantic classes generically and fail to cover the image-specific appearances of target objects, weakening visual-text alignment. To overcome these limitations, we propose olistic mage-adaptive ocabulary volution (), a training-free framework that jointly improves visual representations and textual vocabularies. Specifically, HIVE introduces Holistic Visual Reasoning (HVR) yielding holistic representations through local-global-local interaction. Furthermore, Image-adaptive Vocabulary Evolution (IVE) constructs a dedicated vocabulary set for each image. It first employs an off-the-shelf MLLM to generate image-conditioned category aliases and then iteratively evolutes the vocabulary set by selecting beneficial aliases under semantic and structural visual guidance. Experiments across eight evaluation protocols show that HIVE achieves an average mIoU of 56.2%, outperforming existing training-free OVSS methods and establishing new state-of-the-art performance. Our code will be publicly available after acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.