acceptodds
Under review as a conference paper at ICLR 2027

Forge and Bridge: Strictly Inductive Multimodal Intent Discovery

Abstract

Unsupervised multimodal intent discovery aims to uncover potential intent categories from unlabeled multimodal interactions and generalize the discovered structure to unseen samples. A key challenge lies in the fact that although multimodal large language models can provide rich sample-level semantic descriptions, if we directly rely on MLLMs for classification, it often leads to category collapse. In this paper, we propose a unified framework named IntentForge–IntentBridge, which separates semantic descriptions, intent category discovery, and prediction for unseen samples: IntentForge organizes samples based on their consensus and organizes them into a structured semantic view, from which fixed training partitions are discovered, while IntentBridge migrates this partition to unseen samples through reliability-guided pointwise prediction. The framework uses MLLMs to describe what each sample means, thereby mitigating dominant-category attraction, redundant categories, and fragmented intent structures caused by direct MLLM categorization, and discovers the intent taxonomy from semantic consensus across the entire training set. Our algorithm achieves 62.95%, 47.70%, and 39.49% accuracy on MIntRec, MELD-DA, and IEMOCAP-DA. Especially on the MIntRec dataset, the accuracy rate is far superior to that of existing works, and it also achieves excellent performance on the MELD-DA dataset.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.