acceptodds
Under review as a conference paper at ICLR 2027

GateSlice: Scale-Selective Sliced Fusion for Training-Free Small-Object Detection with Frozen Detectors

Abstract

A fundamental limitation of frozen visual predictors is that potentially useful evidence may become inaccessible when the test input is mapped to the model's fixed operating resolution. For object detection, globally resizing a high-resolution image severely compresses small objects, whereas locally enlarging image regions restores their visual evidence but induces a different prediction distribution because of incomplete context, object truncation, and cross-view confidence shifts. This gives rise to a general inference-time question: how can a frozen model exploit complementary transformed views without allowing less reliable views to destabilize its original predictions? We address this problem through scale-conditioned selective fusion and propose GateSlice, a dual-view, fine-tuning-free framework for small-object detection. GateSlice is based on the observation that view reliability is scale dependent. The global view preserves complete scene context and provides stable predictions, particularly for medium and large objects, while tiled views improve the effective resolution of small objects but are less reliable outside the small-object regime. Accordingly, GateSlice does not treat the two views as equally weighted ensemble members. Instead, it assigns them asymmetric roles: the global branch serves as the prediction anchor, whereas the tiled branch acts as a scale-restricted source of complementary evidence. After tiled detections are mapped to the original image coordinates, an area gate selectively routes only small-scale candidates into fusion. A detector-specific fixed score factor aligns cross-view confidence levels, and class-aware non-maximum suppression resolves duplicate predictions. Together, these designs constrain the influence of the transformed view while preserving useful local evidence. Experiments on COCO, TinyPerson, and VisDrone2019 show that GateSlice consistently improves small-object detection across five frozen detectors, including the transformer-based RT-DETR and Deformable DETR, the one-stage YOLO11n and FCOS, and the two-stage Faster R-CNN. These improvements are obtained without introducing learnable parameters or updating model weights. Multifactor ablations and instance-level transition analyses further show that the gains predominantly arise from recovering small objects missed under global resizing, while scale gating and confidence calibration restrict the interference of tiled predictions with the global results. The consistent gains across distinct detection paradigms indicate that GateSlice is not tied to a particular architecture. Instead, inference-time enhancement of frozen detectors can be formulated as reliability-aware evidence selection across complementary views rather than unconstrained prediction aggregation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.