SCHRONO: CONDITIONAL DIFFUSION MODELING VIA SENSORY-CHRONO HIERARCHY FOR MULTIMODAL MENTAL HEALTH COUNSELING
Abstract
Multimodal counseling integrates visual, acoustic, and textual cues to understand a client’s emotions and provide empathetic support. However, two architectural weaknesses remain. First, the two objectives are coupled: perception updates perturb the distribution of generated text. Second, causal attention restricts each video frame to earlier frames, so a late disclosure cannot revise an early reading. We introduce SCHRONO, a framework for multimodal emotion understanding and counseling response generation. It uses a frozen masked-diffusion language model with two components: the Sensory layer trains encoders, projectors, attention pooling, and light heads for emotion and strategy prediction; the Chrono layer takes video, audio, and dialogue history to generate counseling responses through masked diffusion. Separate trainable pathways decouple perception from generation. Bidirectional attention lets each frame access the full video and dialogue context. On the MESC dataset, SCHRONO outperforms the compared methods in emotion recognition and counseling strategy prediction. Its responses also receive the highest average trustworthiness ratings from LLM judges. Our method also recognizes emotions across datasets and generates coherent supportive responses on ESConv without training on that corpus. Further fine-tuning improves 13.1 points for emotion recognition on IEMOCAP. Replacing bidirectional attention with causal attention in the same model lowers response quality, supporting the value of full-context access.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.