acceptodds
Under review as a conference paper at ICLR 2027

Open-Vocabulary Object Detection for Low-Altitude Scenarios Using RGB-Infrared Data: A Benchmark and A New Method

Abstract

Low-altitude UAV perception plays an important role in applications such as urban management, environmental monitoring, and emergency response. However, conventional object detectors rely on a predefined category set and struggle to recognize emerging object categories in real-world environments. Open-vocabulary object detection (OVD) provides a promising solution by transferring visual-language knowledge to enable recognition beyond fixed categories. Nevertheless, applying OVD to low-altitude scenarios remains challenging due to two critical issues: existing UAV datasets are mainly designed for closed-set detection and lack multimodal, large-scale benchmarks for open-vocabulary learning; meanwhile, current OVD methods are predominantly trained on natural images or remote sensing images, which exhibit substantial domain gaps from low-altitude aerial imagery characterized by small objects, arbitrary orientations, complex backgrounds, and significant illumination variations. To address these challenges, we construct UAVOVD, the first visible-infrared open-vocabulary object detection dataset specifically designed for low-altitude scenarios. UAVOVD contains 25K registered multimodal image pairs, 88 object categories, and 350K rotated bounding box annotations, with a seen-unseen category split to support open-vocabulary generalization evaluation. Based on this dataset, we propose LAOVD, a multimodal open-vocabulary rotated object detection framework for low-altitude environments. Furthermore, we introduce a Structure-Guided Asymmetric Gated Fusion (SGAGF) module, which leverages infrared modality as a stable structural prior and adaptively integrates visible modality details according to cross-modal reliability, thereby improving object localization and semantic transferability under complex imaging conditions. Extensive experiments on UAVOVD, DroneVehicle, and ATR-UMOD demonstrate that the proposed dataset provides a challenging benchmark for low-altitude open-vocabulary detection. Moreover, LAOVD consistently outperforms existing methods in unseen-category detection and cross-domain generalization, validating the effectiveness of multimodal open-vocabulary learning for low-altitude intelligent perception.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.