acceptodds
Under review as a conference paper at ICLR 2027

ReasonAVEdit: Explicit Multimodal Reasoning for Instruction-Guided Audio-Visual Editing

Abstract

Existing audio-visual editing methods face a key trade-off between precise control and a unified instruction interface. Explicitly conditioned methods provide precise control, but they rely on task-specific priors such as spatial masks and target audio. Instruction-guided methods need no extra control, but they lack an explicit multimodal reasoning representation and therefore localize edit targets imprecisely. To resolve this tension, we propose ReasonAVEdit, a Perceive–Reason–Edit framework. ReasonAVEdit requires an audio-visual diffusion model to predict multimodal Reasoning Tokens while generating the target audio-visual tokens. The Reasoning Tokens represent the visual edit region, the acoustic target, and the edit time span shared by both modalities. The model thus forms a complete editing decision on its own, without user-provided masks or target audio. The Temporal Query Bridge (TQB) passes the edit time span from the acoustic Reasoning Tokens to the video Temporal Queries, so the edited lips follow only the target's own speech. The updated queries and a single-frame visual region jointly guide video editing, so video-side reasoning no longer grows multiplicatively with video length and resolution. We build ReasonAVEdit-Data with 200K samples and ReasonAVEdit-Bench with 1,200 samples. Across 33 metrics on our benchmark and two public benchmarks, ReasonAVEdit achieves the best result on 32. Relative to the strongest baseline on each metric, it improves audio-visual alignment by 57.4% and reduces the audio-visual temporal offset by 44.5%. Single-frame visual reasoning also reduces inference latency by 51.4% compared with a full-length mask sequence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.