SetOPD: From Few Visual Exemplars to Multimodal Candidate Sets for Remote-Sensing Open-Prompt Detection
Abstract
Open-prompt detection (OPD) allows users to specify targets with text, visual exemplars, or both. Existing multimodal OPD typically compresses multiple exemplars into a class-level prototype and combines text and vision in prompt or representation space, leaving modality-specific candidate states unavailable for explicit cross-modal comparison. We formulate these two limitations from a set perspective. SetOPD first introduces Base–Residual prompting, which reads boxed exemplars in scene context and maps pooled support evidence into a fixed-capacity anchor–correction visual memory that drives a dedicated visual pathway. We further introduce Paired-Query Arbitration (PQA), in which text and visual readers decode from a shared prompt-conditioned initialization, producing naturally paired candidate states that are arbitrated within each pair and subsequently refined as a set. Across 11 heterogeneous remote-sensing sources, including four fully held-out datasets, PQA outperforms representation-level joint prompting on every source and raises Macro11 AP from 54.00 to 55.57. It also improves over the text and Base–Residual visual readers by 1.66 and 2.02 AP, respectively, while adding only 0.63M trainable parameters. Detection-level analysis further shows that the two readers recover complementary objects and that PQA preserves a substantial fraction of this modality-exclusive evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.