acceptodds
Under review as a conference paper at ICLR 2027

CueMem: Discovering Shared Visual Cues in Long Image Collections

Abstract

Large visual collections contain recurring evidence that links images beyond their dominant content, enabling collection-level organization and exploration. In this paper, we define query-free shared visual cue discovery: given a long image collection, a model must identify image subsets grounded in recurring, interpretable visual cues and describe what each subset shares, without cue-specific queries or a predefined cue inventory. Conventional VLMs must recover sparse recurring evidence amid increasingly many distractors, yet standard token-level supervision does not directly separate supporting images from irrelevant ones. To study this problem, we introduce CUESEL, a diagnostic benchmark with 4k packed collections and 73k human-filtered cue groups spanning object, appearance, instance, and scene cues. To overcome the lack of explicit support modeling, we propose CUEMEM, which transforms internal VLM query/key interactions into cue-induced image affinities and uses inferred supports to ground cue generation. It achieves 0.747 Pairwise Grouping F1 (PGF1), compared with 0.405 for the strongest supervised flat VLM. Its advantage over flat VLMs grows as cue support becomes sparse relative to the collection, indicating that explicit relation learning and support-conditioned generation help VLMs recover recurring evidence under increasing visual distraction. Code and model data will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.