acceptodds
Under review as a conference paper at ICLR 2027

Contextualize to Generate: Hierarchical Semantic Organization for Vision–Language Generation

Abstract

Image captioning translates visual content into natural language and has advanced considerably with visual representation learning and language modeling. However, different generation tasks may rely on different visual evidence. While conventional captions often focus on salient objects and their relationships, specialized reports can require detailed observations distributed across multiple regions. Simply aggregating these observations into a global representation may therefore weaken the local evidence required for generation. To better organize such evidence, here we study how local visual observations and shared context can be used together while preserving the original visual sequence, and propose Hierarchical Semantic Contextualization (HSC). Specifically, HSC first uses learnable semantic anchors to extract semantic information relevant to the task from visual features under paired text guidance. Then, HSC integrates the extracted context with the original visual sequence and leverages Hierarchical Semantic Memory to progressively organize local, structural, and semantic information. In the full model, a variational residual with KL regularization further refines the contextualized representation while preserving its direct path to the generator. As a result, local observations remain available together with information collected across the input. We carefully evaluate the proposed model on whole-slide image (WSI) report generation and MS COCO captioning, covering both specialized report generation and general image captioning. The results show that the proposed model consistently improves the matched baselines, with the full configuration providing additional gains overall. Component and placement analyses further examine the contributions of the contextualization design and residual refinement. These results support the effectiveness of the proposed framework across different generation settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.