Preserve, Reveal, Expand: Towards Faithful 4D Video Editing with Region-Aware Conditioning and Benchmarking
Abstract
While 4D-driven video diffusion models can generate visually plausible videos, faithful 4D editing poses a stricter requirement: source-observed content should be preserved, while disoccluded and out-of-view regions must be synthesized naturally. We identify a key failure mode in existing methods, which we term Evidence-Role Mismatch: source-backed observations, uncertain rendered cues, and unsupported regions are entangled within a single conditioning signal, which can lead to preservation drift, ghosting, and unstable extrapolation. To mitigate this issue, we propose PREX (Preserve, Reveal, Expand), a region-aware framework that decomposes the target spatiotemporal volume into three distinct roles based on observation support and scene extent. PREX traces edited 4D points back to source frames and retrieves observed colors and pairs these appearance cues with geometry-aware confidence and region roles, and injects them into a frozen video diffusion backbone through an adapter trained with self-supervised proxy editing tasks. We further introduce PREBench, a diagnostic benchmark specifically designed for 4D video editing, to expose and quantify such region-specific failures. PREBench separately evaluates Preserve fidelity, Reveal quality, and Expand plausibility of videos generated by 4D driven video diffusion models. Experiments show that PREX substantially reduces region-specific failure modes while maintaining strong visual quality and 4D edit control compared with existing 4D-based video diffusion methods. Code will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.