FIVE: Frequency-Aware and Intent-Guided Visual Evidence Learning for Remote Sensing Composed Image Retrieval
Abstract
Remote Sensing Composed Image Retrieval (RSCIR) is an emerging retrieval paradigm that retrieves target images from complex scenes with multi-scale objects guided by diverse modification intents. However, existing methods still suffer from limitations in both visual representation granularity and modification intent modeling. On the one hand, small objects are prevalent in remote sensing images, yet their visual cues are often overwhelmed by large background regions in the resulting image representations, limiting fine-grained retrieval. On the other hand, most methods inject the query text into the retrieval representation as a unified condition, lacking explicit modeling of the visual evidence selection and semantic transfer mechanisms required by different editing operations. To address these issues, we propose Frequency-aware and Intent-guided Visual Evidence Learning (FIVE), a novel framework for RSCIR that improves the discriminability of retrieval representations from two perspectives: object-level visual representation enhancement and intent-guided semantic selection. Specifically, we design a Frequency-aware Object Enhancement module that introduces the Fast Fourier Transform during visual feature modeling to explicitly capture frequency-domain cues, thereby enhancing the representation of contours, textures, and local structures of small objects. Furthermore, we develop an Intent-guided Visual Evidence Selection module that leverages modification intents to guide the selection and reorganization of object regions, attribute cues, and spatial relations in the reference image, generating intent-aware query representations. Experimental results on two RSCIR datasets demonstrate that our method outperforms existing approaches, and ablation studies further verify the effectiveness of frequency-domain representation enhancement and modification intent modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.