acceptodds
Under review as a conference paper at ICLR 2027

FSD: Bridging Foundation Semantics and Local Details for Real-Time Object Detection

Abstract

Real-time object detection has advanced rapidly with CNN- and Transformer-based architectures. Recently, Vision Foundation Models (VFMs) have been introduced to further enhance detectors with rich pretrained semantics. Where partially overlapping semantic modeling may lead to redundant representations and limit further improvements. In fact, the structural complementarity between the VFM and detector remains largely overlooked in existing VFM-enhanced detection research. Motivated by this, we propose FSD, a VFM-enhanced real-time detector that exploits complementary semantic and spatial representations. Specifically, FSD couples a frozen EuPE semantic branch with a trainable CNN spatial pathway to combine global semantics with fine-grained spatial details. To bridge their heterogeneous representations, we further propose Multi-level Semantic-Detail Interaction (MSDI) for adaptive multi-level interaction. Finally, we design an Efficient Auxiliary Head (EAH) to promote detection-oriented adaptation of semantic features and better balance semantic and spatial-detail cues. Extensive experiments on COCO demonstrate that FSD achieves state-of-the-art performance with a favorable accuracy and efficiency trade-off. Further analyses highlight the importance of representational complementarity between VFMs and detectors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.