EmoTrace-Bench: Diagnosing Emotional Expression in Text-to-Video Generation through Controlled Cues and Intensity
Abstract
Text-to-video (T2V) models are increasingly used to generate videos with targeted emotional expressions, yet existing evaluation methods mainly rely on holistic video–text alignment and provide limited insight into whether failures arise from target-emotion expression, intensity control, or specific affective cues. We introduce EmoTrace-Bench, a fine-grained diagnostic benchmark that controls emotion intensity and selected affective cues in prompts and evaluates generated videos at multiple levels. EmoTrace-Bench contains 448 controlled prompts covering eight emotions, two intensity levels, and diverse cue conditions. Using 12 T2V models, we generate 5,376 videos and evaluate target-emotion presence, perceived intensity, and ten audiovisual dimensions of emotional expression with an automated evaluator calibrated and validated against human annotations on a subset of the data. Results reveal substantial differences across models and prompt designs. Prompts that explicitly describe a character’s emotional reaction are followed more reliably than prompts that describe only events or scenes, whereas adding more affective cues or using more explicit cinematic audiovisual language does not consistently improve emotional generation. These results illustrate how aggregate alignment scores may overlook localized differences in target-emotion expression, intensity control, and cue realization. EmoTrace-Bench provides a traceable and interpretable framework for diagnosing emotional video generation without modifying existing generation pipelines, and we release the prompts, annotation guidelines, evaluation protocols, and automated scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.