acceptodds
Under review as a conference paper at ICLR 2027

Resolution-Aligned Change Detection via Diffusion Super-Resolution and RL-Guided Unseen-Class Adaptation

Abstract

Cross-resolution change detection (CD) is challenging because differences in spatial resolution between temporal observations introduce structural misalignment, inaccurate boundaries, and false change responses. An additional challenge arises when semantic classes excluded from base-detector training may appear after deployment without target-scene pixel-level annotations. We introduce a resolution-aligned adaptation framework that combines conditional diffusion-based super-resolution with reinforcement-learning-guided pseudo-label refinement. First, a conditional diffusion model reconstructs the low-resolution after image in the spatial domain of the high-resolution before image using noise-prediction, reconstruction, and perceptual objectives that preserve geometry and empirically reduce reconstruction-induced artifacts. A shared-weight Siamese detector with cross-temporal attention then predicts changes from the aligned image pair. To adapt beyond the original label space, candidate-mask refinement is formulated as a finite-horizon decision process. An RL policy iteratively refines, accepts, or rejects candidate regions using a retrieval-augmented reward combining geometric consistency, image-text alignment, and language-guided semantic feedback. Accepted masks provide pseudo-labels for target-mask-free adaptation to semantic classes held out from base-detector training. Experiments on synthetic and real cross-resolution benchmarks demonstrate improved reconstruction fidelity, change-detection accuracy, held-out-class coverage, mask quality, and post-adaptation detector performance over matched baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.