DTE-Cap: Learning Directional Temporal Evidence for Remote Sensing Change Captioning
Abstract
Remote Sensing Image Change Captioning (RSICC) aims to understand changes between bi-temporal remote sensing images and describe them in natural language. However, existing methods remain limited in modeling bi-temporal change representations. First, they lack explicit modeling of ordered cross-temporal correspondences, making their predictions dependent on the input order established by the dataset and prone to errors when this order is reversed. Second, changes of different natures are often compressed into a single representation, hindering the language model from accessing structured change evidence with clear semantics. To address these issues, we propose DTE-Cap, a remote sensing change captioning framework for direction-aware temporal relation modeling and structured event representation. Specifically, we design a Directional Temporal Correspondence Modeling module that constructs mutually conditioned semantic references between the two observations, transforming conventional difference representations into relational representations with explicit temporal directionality. Building upon this, we propose a Transport-guided Event Structuring module that formulates cross-temporal semantic organization as an optimal transport problem. It explicitly derives persistent, appearing, and disappearing events from a unified transport coupling and incorporates them into the memory representation to guide caption generation. Extensive experiments and qualitative visualizations on three RSICC benchmarks demonstrate that DTE-Cap consistently outperforms existing methods, producing more accurate and detailed descriptions of visual changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.