acceptodds
Under review as a conference paper at ICLR 2027

DUAL: Closing the Plan–Execution Gap between MLLMs and DiTs for Video Editing

Abstract

Strong multimodal understanding and high-quality video generation do not guarantee faithful video editing: a model may identify the requested change yet edit the wrong object, assign the wrong attribute, or violate the relation that identifies the target. We call this mismatch the plan–execution gap and trace it to the interface between a multimodal large language model (MLLM) and a diffusion transformer (DiT). We propose DUAL, which routes target-state information and relational constraints through different paths. Its Target-State Path converts MLLM features from multiple layers into tokens that join source and noisy target tokens in DiT self-attention, while its Relational-Constraint Path combines final-layer MLLM features with native text embeddings for cross-attention guidance. DUAL reaches the highest Overall score among the compared open methods on both IVEBench (0.67) and OpenVEBench (4.26), while running up to 15.2 times faster than the slowest compared baseline in average inference latency. These results support separating target-state information, which should participate in denoising, from relational constraints, which should guide it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.