MoRE-DETR: Adaptive Recursive Reasoning with Vision-Language Model for Enhancing Remote Sensing Object Detection
Abstract
Remote sensing object detection suffers from severe foreground-background imbalance and substantial variations in object scale and density, making foreground tokens differ greatly in detection difficulty. Existing DETR-like detectors refine queries uniformly and thus under-refine hard objects, while VLM-assisted variants either retain the VLM during inference or generate coordinates as text with weak geometric supervision. To address these issues, we propose MoRE-DETR, a DETR-based framework that couples adaptive token computation with VLM-assisted training on a shared token representation. Specifically, an Adaptive Recursive Token Routing (ARTR) module employs an uncertainty-aware Mixture-of-Recursions router to allocate recursive depth according to token difficulty, producing difficulty-adaptive foreground tokens. These refined tokens are fed to a Visual-Language Aligned Supervision (VLAS) module, which projects them into the semantic space of a multimodal large language model and directly regresses continuous bounding boxes from its hidden states, thereby injecting explicit geometric supervision back into the token representation. VLAS is used only during training and introduces no additional inference overhead. Extensive experiments on multiple remote sensing benchmarks demonstrate state-of-the-art performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.