acceptodds
Under review as a conference paper at ICLR 2027

From Evidence Content to Decision Roles: Role-Conditioned Reasoning for Multimodal Intent Recognition

Abstract

Multimodal intent recognition (MIR) integrates linguistic, acoustic, and visual evidence for inferring a speaker’s intent. Existing methods primarily model what this evidence conveys or how strongly it should contribute to fusion, leaving a more fundamental question underexplored: what role should the evidence play in the current decision? Reliable evidence may be redundant, while conflicting evidence can either correct predictions or introduce noise. Therefore, we propose RCER, a Role-Conditioned Evidence Reasoning framework that organizes multimodal intent inference into an evidence–role–decision pipeline. RCER employs a latent evidence structuring mechanism, which uses shared latent queries to organize heterogeneous modality sequences into slot-aligned, modality-specific evidence and explicitly encodes pairwise relations within each slot. It then performs role-conditioned evidence reasoning to characterize sample-dependent evidence roles through predictive conflict and signed counterfactual utility learned from leave-one-modality-out supervision. RCER encodes these signals as explicit role tokens that condition joint reasoning over unimodal and relational evidence, thereby guiding evidence interactions beyond scalar fusion weighting. Extensive experiments on two challenging MIR benchmarks, MIntRec and MIntRec2.0, demonstrate that RCER achieves state-of-the-art performance, highlighting the value of modeling evidence roles alongside evidence content for MIR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.