acceptodds
Under review as a conference paper at ICLR 2027

ANIMACT: What Breaks When Characters Meet? A Capability Taxonomy for Multi-Reference Animated Video Generation

Abstract

Reference-conditioned animated video generation requires models to preserve visually specified characters and props while executing a short story with the correct actions, interactions, and bindings. Animation makes this evaluation particularly challenging: stylized deformation and exaggerated motion may be intentional, while identity, contact, weight, and who-acted-on-whom must remain visually clear. Existing benchmarks capture parts of this problem, but do not jointly evaluate reference-grounded story execution within a single generated clip. We introduce , a holistic human evaluation benchmark that combines 14 overall criteria with eight targeted capability probes, covering reference fidelity, action execution, role assignment, interaction, identity preservation, and temporal, spatial, and physical stability. The benchmark contains 317 human-verified stories spanning 64 characters and 13 props; 95% are multi-character, and each conditions on 1–4 reference assets together with evaluation-only visual and action annotations. Across 1,902 clips from six reference-to-video systems, Seedance-2.5 is the clear overall leader, yet substantial failures remain: 40% of scripted character-to-character interaction edges do not occur, appearance consistency degrades as reference load increases while entity presence remains stable, and only of clips pass all applicable evaluation axes. Finally, we show that AI judges recover aggregate model trends much better than individual human judgments, making them useful for relative model comparisons but not yet a replacement for fine-grained human evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.