WorldCast: World-Consistent Character Video Generation with Large-Angle Camera Control
Abstract
Precise camera control is essential for cinematic character video generation, yet existing methods struggle to follow large-angle trajectories that combine rotation, translation, and radius variation, with arbitrary initial azimuths. Such viewpoint changes also reveal character regions unseen in the initial frame, often leading to geometric distortions, appearance degradation, and identity drift. To systematically study these coupled challenges, we introduce CLACT, a large-scale Character-centric dataset and benchmark covering diverse Large-Angle Camera Trajectories, comprising 80,000 synthetic training samples and 42 synthetic and real-world evaluation scenes, with accurate camera geometry and multi-view character observations. We further propose WorldCast, a framework that couples scale-aligned geometric guidance with geometry-aligned multi-view character memory. Specifically, we back-project the initial view into a 3D point cloud and render it along the target trajectory to provide trajectory-aligned geometric guidance and coverage masks. To complement character information missing from the initial view, we introduce Geometry-Aligned Memory Attention (GAMA), which aligns multi-view anchor geometry with target-view geometry in a shared coordinate frame, enabling geometry-aware retrieval and injection of view-relevant character features. Experiments across diverse large-angle trajectories demonstrate improved camera-control accuracy and character consistency over the compared methods while maintaining competitive visual quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.