Contextualize to Generate: Hierarchical Semantic Organization for Vision–Language Generation
Abstract
Image captioning translates visual content into natural language and has advanced considerably with visual representation learning and language modeling. However, different generation tasks may rely on different visual evidence. While conventional captions often focus on salient objects and their relationships, specialized reports can require detailed observations distributed across multiple regions. Simply aggregating these observations into a global representation may therefore weaken the local evidence required for generation. To better organize such evidence, here we study how local visual observations and shared context can be used together while preserving the original visual sequence, and propose Hierarchical Semantic Contextualization (HSC). Specifically, HSC first uses learnable semantic anchors to extract semantic information relevant to the task from visual features under paired text guidance. Then, HSC integrates the extracted context with the original visual sequence and leverages Hierarchical Semantic Memory to progressively organize local, structural, and semantic information. In the full model, a variational residual with KL regularization further refines the contextualized representation while preserving its direct path to the generator. As a result, local observations remain available together with information collected across the input. We carefully evaluate the proposed model on whole-slide image (WSI) report generation and MS COCO captioning, covering both specialized report generation and general image captioning. The results show that the proposed model consistently improves the matched baselines, with the full configuration providing additional gains overall. Component and placement analyses further examine the contributions of the contextualization design and residual refinement. These results support the effectiveness of the proposed framework across different generation settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.