acceptodds
Under review as a conference paper at ICLR 2027

Beyond Textual Shortcuts: A Training Framework for Hierarchical Visual Recognition with Vision-Language Models

Abstract

Hierarchical visual recognition requires vision-language models (VLMs) to make accurate and taxonomically consistent predictions across multiple semantic levels. However, existing hierarchical visual instruction tuning methods expose preceding ground-truth answers during training, potentially encouraging textual shortcuts rather than visually grounded recognition. Through a reference-taxonomy analysis and controlled interventions, we reveal that historical textual context can determine many training answers without visual evidence and that trained models exhibit substantial reliance on preceding taxonomic cues. To address this issue, we propose a decoupled hierarchical training framework that separates hierarchical supervision from historical textual dependencies. Specifically, it combines Exchange-Isolated Causal Attention (EICA), which isolates preceding question–answer exchanges while preserving shared visual information, with a probabilistic hierarchy loss that explicitly encourages taxonomic consistency across prediction levels. Experiments on iNat21-Plant and iNat21-Animal with two generative VLM backbones demonstrate consistent improvements in hierarchical consistency and fine-grained recognition across seen, unseen, and cross-domain settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.