InstructCD: A Unified Instruction-Following Framework for Remote Sensing Change Detection
Abstract
Remote-sensing change detection has diverged into four paradigms: binary, semantic, referring, and reasoning, with each addressed by specialized models and increasingly driven by natural-language instructions. Existing unified models cover only part of this spectrum. They merge binary and semantic detection, or couple detection with textual question answering, yet none handle all four paradigms in a single model, and none support variable-category referring with rejection or multi-level reasoning. We present InstructCD, an instruction-following framework that recasts every paradigm as one interface: a language query produces a set of <change> tokens, and each token decodes to a mask, so the number and meaning of the outputs follow the instruction. A Multimodal Language Model conditions the queries, a Temporal Scale Adapter encodes pre-change, post-change, and difference evidence, and a Flexible Mask Decoder emits a variable number of masks. To evaluate the harder paradigms, we build directional semantic queries, a category-referring set, and InstructCD-29K, a reasoning benchmark with four difficulty levels whose supervision is recomputed deterministically from ground truth. Under one protocol in which every baseline uses the same or a larger image encoder, InstructCD achieves the best primary metric on all four paradigms, and a single jointly trained model rivals per-task specialists.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.