Seeing Globally, Detecting Locally: Injecting and Characterising Frozen Vision Transformer Priors in Object Detection
Abstract
Convolutional neural networks (CNNs) provide efficient local inductive biases, while Vision Transformers (ViTs) aggregate global scene context via self-attention. Efforts to unify these complementary strengths into hybrid models often require complex, architecture- or domain-specific coupling, limiting transfer across detector families and obscuring how the two priors interact. To address this, we introduce a lightweight framework that uses a frozen vision foundation model as an auxiliary feature supplier, injecting its final patch representations in one direction into the deepest stage (C5) of a host convolutional backbone. The downstream pipeline remains unchanged, so the host's own detection loss determines how local and global features are mixed. Using DINOv3 as the frozen encoder, we conduct a factorial study of host initialisation and prior availability across one- and two-stage detectors, stratified data regimes on PASCAL VOC and Palm Oil, and full-scale COCO. Across settings, frozen priors consistently accelerate learning. They substitute for and complement conventional pretraining: randomly initialised hybrids surpass pretrained baselines at full PASCAL VOC on both hosts, and pretrained hosts still gain. These benefits depend on host competence and target domain. On PASCAL VOC, pretrained hybrids outperform their baselines by 7.1 AP on YOLOv8s and 13.8 AP on Faster R-CNN at full data, and the Faster R-CNN hybrid matches its full-data baseline with an estimated 11.4× fewer labels. On Palm Oil, this advantage contracts as data increase. We further identify a failure mode in which fine-tuning a COCO-pretrained YOLO detector at its default learning rate erodes accuracy on categories it already detects, with limited data restricting recovery. Overall, this study characterises the learning dynamics that emerge when frozen ViT priors meet trainable CNN representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.