From Evidence Content to Decision Roles: Role-Conditioned Reasoning for Multimodal Intent Recognition
Abstract
Multimodal intent recognition (MIR) integrates linguistic, acoustic, and visual evidence for inferring a speaker’s intent. Existing methods primarily model what this evidence conveys or how strongly it should contribute to fusion, leaving a more fundamental question underexplored: what role should the evidence play in the current decision? Reliable evidence may be redundant, while conflicting evidence can either correct predictions or introduce noise. Therefore, we propose RCER, a Role-Conditioned Evidence Reasoning framework that organizes multimodal intent inference into an evidence–role–decision pipeline. RCER employs a latent evidence structuring mechanism, which uses shared latent queries to organize heterogeneous modality sequences into slot-aligned, modality-specific evidence and explicitly encodes pairwise relations within each slot. It then performs role-conditioned evidence reasoning to characterize sample-dependent evidence roles through predictive conflict and signed counterfactual utility learned from leave-one-modality-out supervision. RCER encodes these signals as explicit role tokens that condition joint reasoning over unimodal and relational evidence, thereby guiding evidence interactions beyond scalar fusion weighting. Extensive experiments on two challenging MIR benchmarks, MIntRec and MIntRec2.0, demonstrate that RCER achieves state-of-the-art performance, highlighting the value of modeling evidence roles alongside evidence content for MIR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.