MoZoo:Bringing Animal Meshes to Life with Video Diffusion
Abstract
Creating cinematic-quality animal effects is highly valuable but remains labor-intensive, as producing realistic animal appearance and motion still relies heavily on manual animation. Recent advances in video generation offer a promising path toward automating this process, yet existing models provide limited structural control over animal motion and geometry. Progress toward such controllable generation is further hindered by the lack of large-scale data pairing animated meshes with realistic RGB videos. We introduce mesh-guided animal appearance synthesis, a new formulation in which an animated mesh specifies geometry and motion, while a text, image, or video reference specifies appearance. To enable learning under this formulation, we construct MoZoo-Data, a 62K-clip synthetic-to-real dataset that combines aligned rendered pairs with filtered pseudo-pairs recovered from real videos via inverse mesh generation. Building on this data, we propose MoZoo, an in-context video diffusion model that synthesizes realistic animal videos while following mesh-defined motion and preserving reference appearance. MoZoo introduces RAR-RoPE to establish correspondence between mesh and target representations while handling temporally asynchronous references, and ADA to incorporate structural and appearance conditions without allowing noisy target features to contaminate the conditioning streams. We further introduce MoZooBench, a 120-case benchmark for evaluating structural consistency and appearance fidelity. Experiments show that MoZoo outperforms baselines trained on the same data and generalizes across reference modalities, geometries, and motion patterns.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.