A-MBER: Affective Memory Benchmark for Multimodal Emotion Recognition
Abstract
Long-horizon interactive systems need to remember not only facts, but also the affective and relational history that gives a user's current utterance its meaning. Existing emotion-recognition benchmarks mainly evaluate local utterance labels or short-context affective cues, while long-term memory benchmarks mainly emphasize factual recall, temporal reasoning, or knowledge update. We introduce A-MBER, an affective memory benchmark for evaluating whether language models can interpret a current anchor turn using historically relevant evidence from prior sessions. A-MBER contains 2,047 benchmark items derived from 34 multi-session interaction trajectories and covers three task families: judgment, evidence retrieval, and explanation. Each item is annotated with content type, memory dependency level, reasoning structure, and robustness condition. We evaluate ten large language models under local-context, full-history, and retrieved-history settings. Results show that historical access substantially improves performance, with the largest gains on long-term implicit affect, high-memory items, multi-hop reasoning, and trajectory-based reasoning. These findings suggest that A-MBER captures a distinct affective-memory capability beyond local emotion recognition and generic long-context question answering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.