acceptodds
Under review as a conference paper at ICLR 2027

AnyDepth-DETR/-YOLO: Any-depth object detection with a single network

Abstract

Most object detectors are static, fixed-depth networks optimized for a single operating point, requiring separate models for different deployment scenarios. We present an any-depth detection framework that enables a single network to span a continuous range of accuracy-efficiency trade-offs by controlling depth at inference time without retraining. Our key design principle is stage-wise modularity: each stage's output remains compatible with downstream stages regardless of depth. To this end, each backbone and neck stage is divided into an essential path, which always executes, and a skippable refinement path, preserving the full multi-scale feature hierarchy at every depth configuration. Training such a network requires jointly optimizing many sub-networks, which introduces conflicting gradient signals. We address this via self-distillation between only the two extremes, with prediction- and feature-level alignment losses that enforce stage-wise modularity. Instantiated on RT-DETR and YOLOv12, our full-depth configurations match or surpass their respective state-of-the-art baselines with negligible parameter overhead, while the most efficient configurations achieve up to speedup at a cost of only 2.0 AP, all from a single set of weights.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.