acceptodds
Under review as a conference paper at ICLR 2027

Small Objects Inside the Grid: Activating Latent Scale Knowledge in Frozen Detectors via Training-Free Grid Reparameterization

Abstract

Small object detection is central to high-resolution visual settings such as aerial perception and long-range surveillance. Its core challenge is not limited to scarce target pixels. It also concerns how continuous changes in input scale preserve a consistent semantic-geometric correspondence on the discrete multi-scale grids of a frozen Transformer detector. This internal scale-adaptation mechanism remains poorly understood. Existing methods commonly rely on retraining, architectural expansion, or black-box multi-scale fusion. These approaches implicitly assume that small-object capability must be reacquired through external supervision or that additional resolution can be directly exploited by the model. We challenge this assumption and posit that a pretrained detector already encodes scale-related positional knowledge, while fixed positional embeddings, decoder anchors, and validity masks constrain its use when the input shape changes. To test and exploit this mechanism, we propose Adaptive Causal Grid Intervention (ACGI), a training-free framework that synchronously reconstructs internal spatial tensors from actual feature shapes. ACGI characterizes the response discrepancy between two internal coordinate systems as cross-grid geometric consistency and uses annotation-free spatial-compression routing to select an intervention scale. It requires no gradients, target-domain annotations, or parameter updates. Experiments on Microsoft Common Objects in Context (COCO), TinyPerson, and VisDrone2019 with multiple frozen DEtection TRansformer (DETR) baselines show consistent improvements in overall and small-object detection across different training strategies and decoder sampling mechanisms. Reliability ranking, scale diagnosis, and deletion analyses further demonstrate that cross-grid consistency faithfully captures the dependence of predictions on internal spatial representations. These results recast scale adaptation in small object detection as the identification and activation of latent spatial knowledge in frozen models, rather than the simple addition of input pixels or new downstream parameters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.