OmniInstruct: Instruction‑Guided Audio-Video Editing via Universal Dataset and Reward Optimization
Abstract
Instruction-guided audio-video editing aims to jointly modify visual and auditory content according to natural-language instructions while preserving source content and cross-modal coherence. While existing methods have achieved promising progress, two challenges still remain: (1) the absence of a large-scale, high-quality dataset covering diverse audio-video editing tasks, and (2) the inadequate generalization ability of editing models on open-domain scenarios. To tackle the first challenge, we introduce OmniInstruct-500K, a high-quality dataset comprising 500K audio-video editing triplets across local, global, and speech editing categories constructed through an automated pipeline. Based on this dataset, we propose OmniInstruct, an end-to-end framework that first adapts a pretrained joint audio-video generator into an instruction-guided editor through task-aware supervised fine-tuning. To tackle the second challenge, we further introduce Edit-NFT, a reinforcement learning post-training method that employs editing-oriented rewards for instruction following, source preservation, perceptual quality, and audio-visual synchronization, and routes reward-wise advantages to the corresponding modality branches. Extensive experiments show our dataset and framework surpass prior methods on quantitative and qualitative metrics, yielding stronger instruction following, editing quality and open-domain generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.