StoryDynamics: Holistic Evaluation of Story Visualization
Abstract
A coherent story visualization requires more than generating individually plausible images. Correct entities should appear, recurring characters should remain recognizable, their interactions should reflect the narrative, visual states should be faithfully depicted, those states should change at the appropriate point in the story, and the resulting frames should remain visually faithful to the text. Existing benchmarks capture subsets of these properties, but do not provide a unified evaluation of them jointly. We introduce StoryDynamics, an evaluation framework that assesses story visualization across six complementary dimensions: entity count, character identity, character-object interaction, visual state, state transition, and visual quality. To support this evaluation, we construct a benchmark of 93 human-validated English and Chinese stories spanning diverse narrative lengths, domains, themes, and visual styles. Each story is paired with the reference information needed to evaluate both frame-level properties and how observable states evolve throughout the narrative. Using StoryDynamics, we evaluate 21 general-purpose and story-specific generation models. Our benchmark evaluation shows that strong performance on individual dimensions does not necessarily translate to coherent story visualization, that character-object interaction is the most vulnerable capability across story length and language, and that story-specific models do not consistently achieve holistic gains over general-purpose text-to-image models. Our results highlight the importance of jointly evaluating both static and temporal properties when assessing story visualization models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.