HumanFX: Human-Centric Visual Effects Generation
Abstract
Visual effects (VFX) combine exaggerated changes in appearance with nonphysical dynamics. Generating these effects for human subjects requires preserving facial identity, body structure, and motion continuity as the effect develops. Multi-person scenes further require accurate effect assignment and coordinated interactions. We present HumanFX, an image-to-video framework for human-centric visual effects generation. First, identity-preserving supervised fine-tuning (SFT) strengthens facial conditioning through a dedicated reference branch. Face crops sampled from video frames provide detailed identity information across varied views and expressions. Second, human-centric reinforcement learning (RL) uses anatomical and motion rewards to improve body structure and temporal consistency, with supervised regularization to retain expressive effect dynamics. Finally, because high-quality multi-person effect data are scarce, we introduce vision-language model (VLM)-guided test-time optimization (TTO). A differentiable visual question-answering (VQA) reward guides updates to the initial noise, refining effect assignment and interaction within a small computational budget. Extensive experiments demonstrate improvements in identity preservation, human motion, effect alignment, and user preference, alongside efficient multi-person control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.