Contextual Cue Replay: Training-Free Whole-Image Emotion Recognition via Progressive Re-reading in MLLMs
Abstract
Multimodal large language models (MLLMs) have made substantial progress in visual understanding, yet whole-image emotion recognition remains challenging, as it requires interpreting visual cues in context and weighing their contributions to the overall emotional impression. Emotion-specific fine-tuning incurs training and data-preparation costs when adapting MLLMs to different emotion-category systems, motivating training-free approaches. While existing training-free methods such as visual token selection help retain relevant evidence, such filtering alone does not establish an explicit feedback path for task-conditioned scene information to refine visual representations. In image-first autoregressive MLLMs, causal attention prevents visual positions from directly accessing such information formed at subsequent instruction positions. We propose **Contextual Cue Replay (CCR)**, a training-free framework that enables progressive contextual re-reading within frozen MLLMs. At selected decoder depths, an auxiliary block evaluation constructs contextualized visual states with access to the complete image and image-conditioned states of a generic task instruction. The base and contextualized states then query a shared memory of the current visual representations. Their retrieval difference captures context-induced changes in retrieved evidence and is written back as an incremental correction to the visual residual. Native decoder blocks integrate each correction before the next replay stage, allowing subsequent re-reading to build on visual evidence and instruction states shaped by earlier updates. Experiments across diverse benchmarks and emotion-category systems demonstrate that CCR yields consistent performance improvements without parameter updates or image-specific human guidance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.