MAniCast: Multi-Shot Animation Creation with Structured Character-Bound Scripts
Abstract
Creating controllable multi-shot animations remains difficult for non-expert users because existing approaches typically require dense frame-level sketches, poses, or keyframes produced with professional drawing and animation skills. A more intuitive alternative is to create animations using only text descriptions and reference images. However, integrating characters defined by text descriptions and reference images with textual narratives remains challenging, particularly in multi-character, multi-shot animation creation. To address this problem, we present MAniCast, the first framework to combine modality-unified character casting with explicit identity-to-role binding for multi-shot animation generation. Structured Character-Bound Scripts combine story- and shot-level instructions with a shared character bank in which each identity is defined by either a detailed text description or a reference image, while explicit placeholders assign selected identities to roles and actions across shots. To realize these assignments at the network level, we introduce Character-Bound Attention-over-Attention (AoA), which derives identity features from textual or visual character definitions and injects them into the corresponding placeholders in the narrative. We further construct MAniScript, a multi-shot animation dataset annotated with Structured Character-Bound Scripts, each of which contains a character bank of text- or image-defined identities. Experiments demonstrate that MAniCast outperforms prior text-controlled and reference-based baselines in casting text- and image-defined characters, executing identity-bound actions, and generating complex interactions across shots. Our model and dataset will be publicly released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.