Stranded Capacity: Deep Pyramid Levels Earn No Return on Small Objects
Abstract
In UAV imagery, remote sensing, and traffic surveillance, small instances account for more than 70% of all objects, against about one third on COCO. Existing work appends modules to a general-purpose detector but inherits the allocation of capacity across pyramid depth, channel width, and input resolution, with the computational budget fixed for a different scale distribution and left largely unexamined. We vary each axis in turn and measure its marginal return per unit of compute. The optimum departs from convention on all three axes, and on two of them the sign reverses: removing deep heads or compressing deep channels raises AP50. We call this condition stranded capacity and trace it to a single geometric quantity, the effective feature area a target occupies on the grid from which the target is predicted. This area scales inversely with the squared pyramid stride and quadratically with input scale. These two dependencies predict the sign of the return on every axis. Reading the configuration from the measured returns yields a three-step transformation that adds no modules. Extensive experiments show that the configuration holds across both detection paradigms. On VisDrone, R-YOLOv11-L matches an identically resolved baseline (41.5 versus 40.8 AP50) with 96% fewer parameters and 76% less compute, and runs 2.3× faster. R-D-FINE-L gains 5.1 AP50 over the corresponding baseline with 87% fewer parameters. The same configuration transfers to five further datasets without retuning, gaining 20.2 and 18.0 AP50 on TinyPerson and SODA-D. The code is included in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.