Steer Unified Audio Editing Toward Reliable Source Selection
Abstract
Unified audio editors support diverse operations, yet often struggle to modify the intended source while preserving non-target content. To strengthen this selective editing capability, we propose Source-selection Training via Early-trajectory Emphasis and Recovery (STEER), a two-stage post-training framework, and use it to develop SteerSound, a unified audio editing model supporting six common tasks. Concretely, intermediate-state observation and bidirectional prefix interventions show that early denoising states causally influence final source selection, while incorrect early decisions often persist through subsequent denoising. Guided by these findings, STEER-Emphasis emphasizes training on task-relevant early-trajectory states to reduce source-selection errors, while STEER-Recovery learns corrective denoising from self-generated rollout states to recover from early-trajectory deviations. We further introduce the Selective Audio Editing Benchmark, enabling large-scale automatic evaluation of target manipulation and non-target preservation. Experiments on both the benchmark and more realistic multi-source audio scenes show that STEER consistently improves selective editing reliability, reducing source-selection errors while improving overall editing performance. Further evaluations across different models and other selective-control tasks also support the board generalizability of STEER and its underlying early-trajectory training rationale. SteerSound achieves advanced performance among open-source unified editors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.