acceptodds
Under review as a conference paper at ICLR 2027

ThermOV: Thermal-Only Open-Vocabulary Semantic Segmentation via Unsupervised Domain Adaptation

Abstract

Open-vocabulary semantic segmentation (OVSS) enables pixel-level recognition of arbitrary visual concepts specified by text. However, existing OVSS methods are predominantly developed for RGB imagery, while prior studies involving thermal data typically rely on co-registered RGB-thermal inputs. As a result, open-vocabulary segmentation from thermal observations alone remains largely unexplored. To bridge this gap, we propose ThermOV, a two-stage unsupervised cross-domain adaptation framework that transfers the open-vocabulary segmentation capability of the RGB-trained foundation model SAM3 to the thermal domain without requiring pixel-level annotations of real thermal images. In the first stage, labeled RGB images are used to generate synthetic thermal images to provide an initial adaptation to thermal appearance. In the second stage, the model is further adapted using unlabeled real thermal images within an EMA teacher-student framework. To preserve open-vocabulary generalization during this process, we construct a diverse prompt bank that excludes categories overlapping with the target evaluation classes and jointly optimize pseudo-label supervision, semantic prompt consistency and teacher-student prediction consistency. Experiments on FMB and TIR-SS show that ThermOV improves over SAM3 by 5.97 and 9.30 mIoU, respectively, and outperforms the strongest RGB OVSS baseline adapted using labeled real thermal data by 3.81 and 0.54 mIoU. These results demonstrate the effectiveness of ThermOV for thermal-only OVSS.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.