HIVE: Hierarchical Induction of Visual Embeddings for Part–Whole Representation Learning
Abstract
Visual recognition benefits from representations that reveal how local image evidence is organized into higher-level semantic structure, yet existing hierarchy-aware visual backbones often provide limited quantitative access to the induced part-whole relations. To address this, we propose HIVE, hierarchical induction of visual embeddings for part-whole representation learning, whose cross-level assignments are directly readable as a part-whole composition graph. HIVE maps image patches to fine-level embeddings, recursively organizes them into higher-level structures through graph-based reasoning and hierarchical pooling, and reintegrates multi-level representations into the token space for recognition. We further introduce part-whole contribution analysis to quantify directional fine-to-mid and mid-to-coarse composition and visualize image-specific parsing trees. Experiments show that HIVE serves as a competitive general-purpose ViT backbone, while producing multilevel embeddings that are class-relevant and structurally coherent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.