acceptodds
Under review as a conference paper at ICLR 2027

Compress Redundancy, Predict What Is Missing: Beyond Global Pooling for Scene Sketch Retrieval

Abstract

Sketch-to-photo retrieval enables users to query images with free-hand drawings, but its central challenge is matching sparse strokes to visually rich photographs. Existing methods emphasize semantic alignment or local matching, yet pay limited attention to two limitations of single-vector aggregation: similar regions are repeatedly encoded, while spatial layout is lost during compression. We argue that learning a compact representation for discriminative retrieval requires consolidating repeated evidence while preserving the spatial ordering of regions. Accordingly, we encode sketches and photos with a shared pretrained DINOv3 ConvNeXt-B backbone and introduce redundancy-aware balanced semantic compression and masked spatial prediction to learn complementary global, regional, and spatial representations. Under identical protocols on FS-COCO, SketchyCOCO-SL, and QMUL-ShoeV2, our method improves upon the previous best results by 8.6%–32.7% in R@1, 5.9%–10.6% in R@5, and 1.9%–6.3% in R@10 across scene-level and fine-grained object retrieval. This provides a new direction for compact representation learning in sketch-driven cross-modal image retrieval, jointly suppressing redundancy while preserving spatial structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.