One Image, Parallel Decisions: Structured Querying for Zero-Shot Multi-Label Recognition
Abstract
Contrastive vision–language models such as CLIP have been widely adopted for zero-shot multi-label recognition, while generative multimodal large language models (MLLMs) have received comparatively less attention in this setting. Applying MLLMs to this task requires an interface that produces scores for multiple non-exclusive categories from generative predictions. Direct label generation does not readily provide a score for every candidate class, whereas evaluating class-specific questions separately repeats shared context computation. We propose Structured Parallel Querying (SPQ), a training-free method that organizes class-presence questions as independent suffixes over a shared visual and instructional prefix. Block-isolated attention prevents information exchange between class suffixes, while aligned positional encodings preserve the context of standalone queries. SPQ reads Yes/No logits at all class-specific decision positions in a single joint forward pass, producing multi-label scores without autoregressive answer generation. Experiments with four MLLM configurations spanning Qwen2.5-VL, LLaVA, and InternVL demonstrate the applicability of SPQ across model families and scales. A direct comparison with independent binary querying further shows that joint evaluation can preserve recognition performance. These findings support structured parallel querying as a practical approach to zero-shot multi-label recognition without task-specific training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.