Four Cells Are Not the Same Dose: Size-Conditioned Spatial Lesions in Detection Transformers
Abstract
A spatial lesion that removes four feature cells can erase most of a small object's footprint while barely affecting a large one. We test whether this mismatch changes conclusions drawn from decoder cross-attention in frozen DETR and Conditional DETR checkpoints. On a preregistered 4,000-image COCO holdout, the small-minus-large local-quality contrasts for four in-box versus four out-of-box cells are 0.02136 and 0.02325. When we instead remove the same 0.02 clean mean-head attention mass from each eligible query, both 95% image-bootstrap intervals fall inside the prespecified ±0.005 equivalence region. This low-dose result applies only to eligible GT-matched objects: the four-cell gate excludes 60.3% of candidate small objects. At an exploratory dose of 0.40, size contrasts rise again, so equal mass is not a general explanation of size sensitivity. Prediction-selected detector AP effects are small and differ from the GT-matched local estimand. An exploratory deformable sampling-footprint candidate improves small-object AP in an author-designated summary, but the baseline fails a prespecified feasibility gate and conflicting source records preclude a validated improvement claim. The practical contribution is a dose-aware spatial-ablation protocol that reports eligibility and the detector-level readout alongside query-level effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.