acceptodds
Under review as a conference paper at ICLR 2027

PSYCHE: Modeling Conversational Emotion as a Continuous Trajectory via Causal Source Decomposition

Abstract

Multimodal Emotion Recognition in Conversation (MERC) faces a fundamental challenge: textual, acoustic, and visual expressions are sparse, regulation-filtered observations of internal emotional states that evolve continuously. Genuine calm and suppressed disappointment can exhibit nearly identical surface signals yet arise from opposite trajectories. Existing methods harness rich conversational context and multimodal cues, yet model emotion as a discrete per-utterance label rather than an evolving state. Shifts are consequently detected by auxiliary binary heads, while causes are attributed post hoc by separate modules, so that how emotion evolves and why it changes are never decoded by a single process. We propose PSYCHE, which models each speaker's emotion as a latent trajectory governed by a Neural Jump ODE: Causal Source Decomposition (CSD) structures the driving force for continuous inter-utterance evolution, while speech acts trigger discrete updates. CSD decomposes this driving force into a self-cause driver with state-dependent decay and an other-cause driver attending to the interlocutor's triggering utterances, fused by a per-dimension gate. The self-/other-cause balance and trigger identification are computed by the forward pass itself, requiring no post-hoc explainer or inference-time language model. On IEMOCAP and MELD, PSYCHE achieves the best weighted F1 among the baselines we re-evaluate under identical multimodal features, with offline commonsense cues as an added input. Dynamics analyses show that the model lowers the self-cause ratio when emotion turns and reduces confusion precisely where surface signals mask internal states, indicating that how and why are not two answers but one mechanism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.