acceptodds
Under review as a conference paper at ICLR 2027

Knowing Whom to Erase: Instruction-Routed Audio-Visual Subject Removal for Human-Centric Speech and Music Videos

Abstract

Audio-visual subject removal erases the person specified by an instruction and the sound they produce, while preserving other people and their sounds. Existing audio-visual editors learn text- or mask-conditioned edits from paired examples, but matching individual edits does not establish reliable target switching within the same recording. In multi-person scenarios, existing editors struggle to jointly remove the requested person and preserve competing sources. Single-target supervision can contribute to this failure by permitting an instruction-agnostic shortcut: when every instruction for an input refers to the same target, a fixed edit can fit the training data without learning whom to remove. To address this problem, we propose Audio-Visual Instruction-Routed Erasure (AV-IRE), a joint audio-video editor combining complementary-target supervision with explicit target routing. Source recordings and their teacher-edited counterparts are composed to supervise different target-specific outputs for the same input. The target is conveyed to the generator through a spatial routing map and modality-specific codes, while its text prompt remains target-neutral. Route-dependent output changes are supervised by a shared-state route-intervention loss. We also construct 54,705 automatically filtered, teacher-generated removal pairs from real recordings for task adaptation. AV-IRE outperforms strong audio-visual editing baselines in selective multi-person removal. Compared with Task-Adapted LTX-2.5, blind-rated joint success rises from 11.7% to 41.7% on natural two-musician recordings. On new recordings, stricter bidirectional success, requiring both target requests per input to succeed, rises from 3.3% to 26.7%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.