acceptodds
Under review as a conference paper at ICLR 2027

Structure Shapes Semantics: A Unified Framework for Text-Assisted Image Clustering

Abstract

Text-assisted Image Clustering (TIC) leverages vision-language models to provide external textual semantic information, thereby enhancing clustering performance. Existing approaches typically follow a two-stage paradigm: semantic representation construction followed by clustering network training. However, these methods heavily rely on handcrafted strategies or point-to-point matching for semantic representation construction, making them vulnerable to imperfect image-text alignment and semantic noise. In this work, we first introduce a principled framework that reformulates semantic representation construction as a global optimization problem to address these challenges, and propose to solve it via coordinate descent. Then, guided by the principle that "tructure hapes emantics" (), we introduce several semantic-quality criteria and formulate them as optimization objectives. Building upon this reformulation, we further unify the two TIC stages under the shared objective and reveal their symmetric optimization processes. To further refine the semantic representations, we propose the Fine-Structure guided Optimization (FSO). Extensive experiments validate the effectiveness of our method, with a 6.0% accuracy gain on the most challenging Tiny-ImageNet. We believe this work advances TIC from empirically driven heuristics toward a principled optimization paradigm. Code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.