TEB: A Bandwidth-Constrained Evidence Allocation Framework for Enhancing Multimodal Sentiment Analysis and Emotion Recognition in Large Language Models
Abstract
Adapter-based LLM frameworks for multimodal sentiment analysis and emotion recognition in AI conversations are attractive for their lightweight flexibility, but the adapter interface provides only tokens. Long audio-visual streams must be compressed into a few evidence tokens while preserving conversation-relevant cues. The key challenge is deciding which evidence to retain under this bandwidth constraint. We propose the Text-Conditioned Evidence Bottleneck (TEB), a bandwidth-constrained evidence allocation framework. TEB uses transcript embeddings to build utterance-relevant query states, organizes audio-visual sequences into local segment and latent global memories, and routes evidence per query through soft gating based on audio-text and visual-text consistency discrepancies. Low-rank coefficients and a shared prompt dictionary map each query to one LLM-receivable evidence token. On three LLMs and four public datasets, TEB improves the five-seed mean of all metrics over MSE-Adapter. At , it reduces ChatGLM3-6B's MAE on SIMS-V2 from 0.279 to 0.265 and increases weighted F1 on CHERMA from 72.80 to 74.18. LLaMA2-7B shows significant improvements in Acc-2, F1, weak sentiment accuracy, MAE, and correlation coefficient on SIMS-V2, while Qwen-1.8B also demonstrates consistent gains on MOSEI and CHERMA. These results establish TEB as a lightweight, flexible, and scalable foundation to advance LLMs' multimodal sentiment analysis and emotion recognition capabilities. Our code is available at https://anonymous.4open.science/r/TEBTEB-111.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.