acceptodds
Under review as a conference paper at ICLR 2027

EmoPrune: Learning Where to Look for Multimodal Emotion Recognition

Abstract

Multimodal large language models (MLLMs) encode each image into hundreds of visual tokens, which make up most of the input sequence and thus most of the inference cost in multimodal emotion recognition (MER). Token pruning can reduce this cost, but for MER it is difficult to decide which tokens to keep: unlike visual question answering or grounding, an emotion prompt does not indicate where the evidence lies, and emotional cues are often small and not visually prominent. Generic pruning signals, which mainly reflect visual prominence, therefore tend to discard them. We propose , an emotion-aware visual token pruning method that learns which tokens to keep. A lightweight Evidence Head, trained offline from the frozen MLLM's answer loss with only image-level labels, scores visual tokens, and its scores are combined with [CLS] attention to retain both local cues and broader visual context. Unselected tokens are removed before the language model, without predefined regions or additional inference stages. Experiments on six benchmarks and two MLLM backbones show a favorable accuracy–compute trade-off. On LLaVA-1.5-7B, retaining 10% of visual tokens reduces computation from 8.8 to 1.78 TFLOPs per image while keeping accuracy close to full-token inference (% vs. 47.81%).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.