DUAL: Closing the Plan–Execution Gap between MLLMs and DiTs for Video Editing
Abstract
Strong multimodal understanding and high-quality video generation do not guarantee faithful video editing: a model may identify the requested change yet edit the wrong object, assign the wrong attribute, or violate the relation that identifies the target. We call this mismatch the plan–execution gap and trace it to the interface between a multimodal large language model (MLLM) and a diffusion transformer (DiT). We propose DUAL, which routes target-state information and relational constraints through different paths. Its Target-State Path converts MLLM features from multiple layers into tokens that join source and noisy target tokens in DiT self-attention, while its Relational-Constraint Path combines final-layer MLLM features with native text embeddings for cross-attention guidance. DUAL reaches the highest Overall score among the compared open methods on both IVEBench (0.67) and OpenVEBench (4.26), while running up to 15.2 times faster than the slowest compared baseline in average inference latency. These results support separating target-state information, which should participate in denoising, from relational constraints, which should guide it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.