ThermOV: Thermal-Only Open-Vocabulary Semantic Segmentation via Unsupervised Domain Adaptation
Abstract
Open-vocabulary semantic segmentation (OVSS) enables pixel-level recognition of arbitrary visual concepts specified by text. However, existing OVSS methods are predominantly developed for RGB imagery, while prior studies involving thermal data typically rely on co-registered RGB-thermal inputs. As a result, open-vocabulary segmentation from thermal observations alone remains largely unexplored. To bridge this gap, we propose ThermOV, a two-stage unsupervised cross-domain adaptation framework that transfers the open-vocabulary segmentation capability of the RGB-trained foundation model SAM3 to the thermal domain without requiring pixel-level annotations of real thermal images. In the first stage, labeled RGB images are used to generate synthetic thermal images to provide an initial adaptation to thermal appearance. In the second stage, the model is further adapted using unlabeled real thermal images within an EMA teacher-student framework. To preserve open-vocabulary generalization during this process, we construct a diverse prompt bank that excludes categories overlapping with the target evaluation classes and jointly optimize pseudo-label supervision, semantic prompt consistency and teacher-student prediction consistency. Experiments on FMB and TIR-SS show that ThermOV improves over SAM3 by 5.97 and 9.30 mIoU, respectively, and outperforms the strongest RGB OVSS baseline adapted using labeled real thermal data by 3.81 and 0.54 mIoU. These results demonstrate the effectiveness of ThermOV for thermal-only OVSS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.