Bayesian Reasoning over Latent Editing Programs for Instruction-Based Video Editing
Abstract
Instruction-based video editing requires a model to determine what to change, where to apply the change across space and time, and what to preserve. Existing Diffusion Transformer (DiT) editors learn these decisions only implicitly, providing the same conditioning to every block and supervising attention solely through the generation objective. We propose **ReVEDiT** (**Re**asoning-based **V**ideo **E**diting with **Di**ffusion **T**ransformers), which formulates video editing as hierarchical Bayesian inference over a latent editing program. First, a routing posterior determines the depth at which native textual and visual tokens enter the DiT, so that shallow blocks form the editing intent from MLLM-derived primitives and deeper blocks ground it in the source video. Second, the ideal attention of the edit is treated as a latent variable, for which a reference branch conditioned on the ground-truth edit provides privileged evidence during training; the resulting posterior is distilled into the editing branch at the grounded layers. At inference, the split layer is fixed and the reference branch is discarded, so editing requires a single forward pass. On OpenVE-Bench and EditVerseBench, ReVEDiT achieves the best overall performance among open-source methods under two MLLM judges, with particularly strong results on localized edits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.