acceptodds
Under review as a conference paper at ICLR 2027

Diagnose Before You Design: Depth–Scale Dependencies in Foundation-ViT Detectors

Abstract

Detectors built on plain foundation Vision Transformers (ViTs) must construct multiscale spatial memories from single-resolution backbone features, yet existing interfaces typically predefine how backbone depth is mapped to detector scale without directly measuring the resulting functional dependencies. Rather than repeatedly retraining alternative routing structures, we recast interface analysis as an interventional diagnosis problem: controlled inference-time interventions on a reference detector quantify the functional dependence of existing depth-to-scale paths and spatial memories. Across DINOv2, DINOv3, and distinct detector interfaces, we consistently observe strong and image-specific depth–scale dependencies, while their dominant spatial realization changes with backbone representation and detector topology. Moreover, weak marginal dependence in a fixed model does not imply retraining-time structural redundancy. Controlled factorial experiments reveal that the effect of depth treatment is modulated by spatial memory topology, establishing the two as coupled design variables and motivating a diagnosis-guided redesign workflow. Guided by these measurements, we formulate and validate distinct interface redesign hypotheses across DINOv3- and DINOv2-based detectors that preserve aggregate Full-COCO accuracy while reducing model capacity, demonstrating the practical value of diagnosis-guided interface design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.