The Change You Want To Talk About: Enabling spatial and narrative Change Detection in Earth Observation with a single VLM architecture
Abstract
Bi-temporal change detection (CD) of Earth Observation images is fragmented into disconnected spatial and textual prediction tasks. While recent Vision Language Models (VLM) explore open-vocabulary and query-driven challenges, their semantic flexibility is often constrained by decoupled, post-hoc matching pipelines. Besides, they often struggle to capture fine-grained temporal dynamics and remain restricted to isolated tasks, such as dense prediction, captioning, or grounding. We propose LOSCAR, a multi-task VLM capable of addressing diverse CD tasks under a unique paradigm, through specific prompting. By integrating a spatio-temporal architecture with a spatial decoder that translates multimodal features into dense image masks, our model jointly generates natural language descriptions and explicit spatial outputs under a continuous open-vocabulary regime. After multi-dataset training, our experiments show that LOSCAR achieves state-of-the-art results in change captioning while delivering accurate spatial localization and robust reasoning on complex semantic change trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.