Beyond Feature Geometry: Semantic Evidence for Contrastive Multi-View Clustering
Abstract
Multi-view clustering (MVC) aims to discover semantically meaningful groups by exploiting consistency and complementarity across views. Many contrastive MVC methods construct supervision from learned representations or model predictions. However, feature geometry may capture attributes that do not align with the task-specific semantic criterion. Consequently, even high-confidence supervision can reinforce semantically incorrect relationships. We refer to this mismatch as Semantic-Blind Correspondence (SBC). To address SBC, we propose Semantic Evidence for Contrastive Multi-View Clustering (SECMVC), which introduces semantic evidence through a small budget of multimodal large language model (MLLM) queries. SECMVC first constructs verified semantic anchors under a fixed task criterion and partitions the remaining samples into reliable and unreliable sets. Reliable samples learn directly from semantic prototypes without additional queries. For unreliable samples, SECMVC queries a representative subset and uses confirmed assignments to guide reference-based contrastive learning, while uncertain responses provide no semantic supervision. In this way, SECMVC propagates semantic evidence when feature-based assignments are reliable and directs MLLM queries toward uncertain cases. Experiments on four multimodal clustering benchmarks demonstrate the effectiveness of SECMVC.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.