WorldPortal: Recast Residual Correction for Instruction-Guided Video Background Replacement and Relighting
Abstract
Instruction-guided video background replacement and relighting require a model to synthesize a new environment and plausible illumination while preserving the identity, structure, texture, and motion of the foreground subject. Existing methods struggle to balance these objectives: training-free pipelines rely on rigid structural conditions that restrict natural lighting changes, whereas directly trained models degrade foreground structures and details. We present WorldPortal, a unified video editing framework based on Recast Residual Correction. During inference, a forward branch first predicts the video under the target background and illumination. A backward recasting branch then places the edited foreground state back into the source background domain, where its discrepancy from the source trajectory primarily captures foreground discrepancies introduced by editing. Transferring this residual to the target-domain prediction restores foreground details. To support both branches, we further introduce a joint VLM-DiT architecture that adaptively parses the instruction and input video into semantic-, foreground-, and background-centric token groups, providing flexible guidance. Experiments demonstrate that WorldPortal achieves superior overall performance in instruction alignment, lighting quality, structural fidelity, and fine-detail preservation, while introducing only modest additional inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.