CPCNet: Cross-Depth Pyramid Construction for Underwater Object Detection
Abstract
Adapting pretrained vision transformers to underwater object detection requires constructing a multiscale pyramid from encoder features that share a common spatial grid. How features from different encoder depths are combined during this construction determines the representations supplied to the detector. We propose CPCNet, an underwater object detector built around a Cross-depth Pyramid Constructor (CPC). CPC jointly aggregates four projected DINOv3 states on their native grid to form a shared central representation. This center feeds a finer branch and a progressively coarser chain, with depth-specific bypasses providing direct encoder inputs to the branches. In single-run evaluations, CPCNet achieves 69.66, 65.27, and 52.18 AP on DUO, RUOD, and UTDAC, respectively. Across all three underwater datasets, CPC achieves higher AP than approximately parameter-matched independent routing and higher AP with fewer operations than source-complete parallel construction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.