AffectVBench: Benchmarking Affective Faithfulness in Long-Form Text-to-Video Generation
Abstract
Existing benchmarks have accelerated long-form text-to-video (T2V) generation by assessing visual quality, temporal coherence, and cinematic presentation, but overlook affective faithfulness: whether the intended affect is correctly bound to characters and events, develops coherently across shots, and is supported by perceivable evidence. We introduce AffectVBench, a fine-grained benchmark for this capability. Grounded in Story Prompts and structured references capturing open-vocabulary affective states, intensity, mixtures, and temporal evolution, AffectVBench provides nine evidence-grounded metrics across Accuracy, Development, and Support. It further incorporates blinded anchor-based calibration for Viewer Affect and multidimensional uncertainty-aware probabilistic scoring. Experiments on 13 representative T2V models reveal that current models struggle with multi-character emotion binding, cross-shot affect persistence, and event-driven affect transitions. These findings establish affective faithfulness as a distinct bottleneck for credible long-form video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.