acceptodds
Under review as a conference paper at ICLR 2027

Unlocking All-in-one Restoration From Pre-trained Diffusion Models

Abstract

Pre-trained text-to-image and text-to-video diffusion models (PTDMs) have enabled significant advancements in restoration by serving as powerful clean-image priors. Existing restoration methods typically exploit these priors through backbone fine-tuning, ControlNet-style modules, explicit degradation operators, or inversion. In this work, we investigate whether a frozen PTDM intrinsically possesses restoration capabilities that can be elicited through its native conditioning pathway to directly steer noisy degraded states toward clean outputs. We first probe the conditioning pathway of PTDMs through text prompts and text-token optimization, and find that neither effectively enables restoration. Interestingly, optimizing the conditioning prompt directly at the text-encoder output unlocks strong restoration behavior from the PTDM. However, naively learning this prompt to denoise noisy degraded states to a clean output is unstable due to a misalignment with the reverse sampling trajectory. To resolve this, we learn prompts within a diffusion bridge formulation that aligns training and inference dynamics, enforcing a coherent denoising path from noisy degraded states to clean images. Furthermore, we introduce a sampling-time partial inversion strategy to improve restoration performance and a prompt composition scheme for handling mixed degradations. Instantiated on the WAN text-to-video and FLUX text-to-image models, our approach elegantly enables these PTDMs to achieve strong and competitive restoration performance. The code will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.