acceptodds
Under review as a conference paper at ICLR 2027

From Composites to Control: Geometry-Grounded Supervision for Visual Effects Generation

Abstract

Recent visual effects (VFX) generation methods primarily rely on large video generative models to synthesize both effects and their interactions with scenes. While this paradigm enables flexible creation, their reliance on uncertain generative priors makes effect fidelity and geometric controllability difficult to ensure, especially under camera or subject motion. We introduce effect-faithful geometry-grounded VFX generation, a formulation that decouples reliable VFX supervision from generative uncertainty. Instead of using a video generator to synthesize VFX videos, we composite curated VFX assets onto real-world footage using recovered camera motion, scene geometry, and human motion. This preserves natural video dynamics and asset-defined effect appearance while providing explicit geometric supervision. Based on this principle, we build VFXShot, a dataset of 10,298 videos with effect semantics, camera trajectories, and spatial anchor annotations. We further propose VFXDirector, a unified video diffusion framework that conditions on effect semantics, camera trajectories, and spatial attachment. Experiments demonstrate that effect-faithful geometry-grounded supervision substantially improves controllable VFX generation, enabling VFXDirector to achieve improved video quality, camera consistency, and spatial attachment. These results suggest that reliable compositional supervision provides an effective alternative to purely generative data construction for learning controllable visual effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.