acceptodds
Under review as a conference paper at ICLR 2027

SVEET: Public-Data Design and Multi-Scale Flow Matching for Singing Voice Enhancement

Abstract

Singing voice enhancement (SVE) restores studio-quality vocals from degraded recordings, but clean singing data are scarce and existing systems often rely on proprietary stems or speech models adapted with limited singing data. We study SVE from a public-data design perspective, systematically varying the composition and scale of public speech and singing corpora while constructing a controlled benchmark for both clean-input preservation and restoration. Building on this data design, we introduce SVEET, a conditional flow-matching enhancer, and SVEET-MS, which adds a degradation-adaptive multi-scale objective for spectrally coherent restoration. Both outperform public zero-shot baselines and a strong discriminative system on singing enhancement, while transferring to speech enhancement with the same checkpoint. On the real-world SingVERSE benchmark, they achieve, to our knowledge, the best perceptual quality among systems with publicly available weights. In a blind listening test, SVEET-MS is preferred over SVEET in 71-76% of decided trials despite only modest differences in objective metrics. Code, data recipes, evaluation protocols, and checkpoints are released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.