acceptodds
Under review as a conference paper at ICLR 2027

CONCORD-SAM3: Native Evidence Coordination for Training-Free Open-Vocabulary Remote Sensing Semantic Segmentation

Abstract

Open-vocabulary remote sensing semantic segmentation aims to assign pixel-level labels to arbitrary textual categories beyond predefined land-cover taxonomies. However, independent prompt responses lack explicit cross-class coordination, while textual descriptions alone do not provide image-specific visual references for diverse within-class appearances. We introduce CONCORD-SAM3, a training-free framework built on frozen SAM 3 that coordinates native evidence across competing classes and within-image appearances. Dual-Head Evidence Coordination (DHEC) calibrates class scores using shared visual–text descriptors and accumulates pixelwise instance support through a bounded noisy-OR rule. Positive semantic–instance soft-union residuals are gated by query presence and semantic support within native masks. Multi-Prototype Visual Calibration (MPVC) constructs multiple prototypes per class from the current image's multi-scale features. A prediction-margin gate strengthens visual calibration at ambiguous pixels while reducing intervention where scores are well separated. The framework uses one image encoding per image or crop and one prompt evaluation per supplied query, without target-domain annotations, parameter updates, auxiliary backbones, re-querying, or transformed views. Evaluations on 16 datasets demonstrate the effectiveness of our method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.