Semi-Supervised Multimodal New Intent Discovery via Progressive Learning and MLLM-Guided Semantic Correction
Abstract
Semi-supervised new intent discovery aims to identify unseen intents using labeled examples of known intents and unlabeled data. Existing methods primarily rely on text and overlook complementary acoustic and visual information. Moreover, semantic feedback from large language models can preserve incorrect intent groupings when conveyed through summaries of the current clusters or local rewriting. We propose PLSC, a method that combines progressive multimodal representation learning with semantic correction guided by a multimodal large language model (MLLM). To mitigate instability in semi-supervised multimodal clustering when modality cues are inconsistent, the progressive strategy stabilizes shared multimodal representations through supervised contrastive learning using known intent labels before each clustering round and gradually increases the contribution of pseudo-label classification. To correct erroneous intent groupings, we incorporate local semantic relations inferred by the MLLM from representatives as soft constraints in centroid refinement, allowing the corrected centroids to guide subsequent assignments and optimize intent boundaries. Experiments on MIntRec, MIntRec2.0, and MELD-DA demonstrate consistent gains of approximately 1% to 3% over existing methods across different known class ratios, providing a methodological foundation for semi-supervised multimodal clustering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.