acceptodds
Under review as a conference paper at ICLR 2027

CrossFreq-UHR: Context-Aligned Frequency Learning for Ultra-High-Resolution Object Detection

Abstract

Sparse global–local detectors improve ultra-high-resolution object detection by allocating high-resolution computation to a limited number of informative regions. However, their global observations are mainly used for region selection and downstream decoding, leaving selected patches to form their multi-scale representations without location-specific scene context before aggregation. We refer to this limitation as the pre-aggregation representation bottleneck. To address it, we propose CrossFreq-UHR, a context-aware frequency learning framework for sparse UHR detection. CrossFreq-UHR maps global features to the coordinate system of each selected patch to provide spatially aligned scene context, and decomposes local features into a smoothed structural representation and a detailoriented residual. An asymmetric interaction mechanism conditions the structural path on aligned global context while independently refining patch-specific detail, introducing scene awareness without directly perturbing local fine-scale information. The enhanced representations are subsequently aggregated for query-based global–local detection. Experiments on SODA-A and STAR demonstrate that the proposed pre-aggregation conditioning consistently improves detection accuracy over the sparse global–local baseline with modest additional computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.