CinePersona: Camera-Aware Dynamic Persona Encoding for Identity-Preserving Human Video Generation
Abstract
Cinematic human video generation must jointly control who appears in frame, what the subject does, and how the camera moves, yet the literature has split into single-axis specialists whose serial composition compounds error on every axis. We argue that the right conditioning object is a chronologically ordered space-time lattice of micro-clips cut from a reference video, and present CinePersona, a camera-aware reference encoder for a frozen 14B video diffusion transformer. Its Chrono-Mosaic Lattice packs nine time-ordered micro-clips into one trunk-compatible tensor, and we show that this ordering is what keeps camera direction decodable: time-reversed trajectories are provably indistinguishable under any shuffled interface. A dual-pathway encoder pairs a diffusion-synchronized implicit stream with explicit persona, action and camera branches, Cine-RoPE writes the target trajectory into the trunk's rotary phases, and a tri-axis sampler exposes one guidance dial per control, all while training 1.5% of parameters. On CineHuman-Bench and OpenS2V-Eval, CinePersona natively controls all three axes and achieves state-of-the-art persona similarity, action alignment and camera fidelity, the latter verified with an estimator decoupled from training, on rendered scenes with ground truth, and by 41 human raters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.