acceptodds
Under review as a conference paper at ICLR 2027

Restora: Controlling Generative Freedom for Faithful and Universal Audio Restoration

Abstract

Universal audio restoration aims to recover speech, singing, and music from diverse and compounded degradations with a single model. Generative models suit this goal because they synthesize clean content rather than simply filtering the degraded input, but this creative freedom can also produce hallucinations: plausible-sounding output unsupported by the original signal, so controlling how far generation departs from the input is essential. Two standard mechanisms govern this freedom: reward-based post-training refines the model toward perceptual quality targets, and inference-time guidance steers each sample toward the observed condition. As commonly implemented, each has a specific weakness that we address. Uniform reward optimization pushes every audio domain equally hard, so a domain that has saturated continues to receive noisy updates that degrade held-out quality; Headroom-Anchored RL instead tracks each domain on its own and, once reward signals become unreliable, shifts that domain back toward a stable supervised target. Classifier-free guidance (CFG) steers away from a null reference that discards the degraded input, the very evidence that faithful restoration relies on; Condition-Swap Guidance (CSG) replaces that null reference with a permuted copy of the input's own conditioning segments—same content, wrong timing—giving a structurally aware direction that improves perceptual quality and content accuracy together. Combining both mechanisms, Restora, a single 500M flow-matching model, raises the mean opinion score of degraded inputs from 2.55 to 4.12 across six test sets and achieves the best spectral-fidelity and aesthetic scores among compared music systems. CSG also transfers to a discrete-token speech enhancement model, showing it is not limited to diffusion guidance. Demo page: https://restora-demo.netlify.app/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.