Beyond Textual Shortcuts: A Training Framework for Hierarchical Visual Recognition with Vision-Language Models
Abstract
Hierarchical visual recognition requires vision-language models (VLMs) to make accurate and taxonomically consistent predictions across multiple semantic levels. However, existing hierarchical visual instruction tuning methods expose preceding ground-truth answers during training, potentially encouraging textual shortcuts rather than visually grounded recognition. Through a reference-taxonomy analysis and controlled interventions, we reveal that historical textual context can determine many training answers without visual evidence and that trained models exhibit substantial reliance on preceding taxonomic cues. To address this issue, we propose a decoupled hierarchical training framework that separates hierarchical supervision from historical textual dependencies. Specifically, it combines Exchange-Isolated Causal Attention (EICA), which isolates preceding question–answer exchanges while preserving shared visual information, with a probabilistic hierarchy loss that explicitly encourages taxonomic consistency across prediction levels. Experiments on iNat21-Plant and iNat21-Animal with two generative VLM backbones demonstrate consistent improvements in hierarchical consistency and fine-grained recognition across seen, unseen, and cross-domain settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.