Distill the Prior, Drop the Diffusion: Compact Feed-Forward Text Restoration from Diffusion Priors
Abstract
Diffusion models provide powerful generative priors for image restoration, but transferring such knowledge beyond the diffusion architecture remains challenging. Can a restoration prior learned by a large diffusion model be absorbed by a fundamentally different feed-forward network, despite the mismatch in internal representations and inference mechanisms? Moreover, under aggressive model compression, can this prior still be retained rather than collapsing with student capacity? We study these questions in text image super-resolution and introduce PixelSTR, a diffusion-prior distillation framework that transfers restoration behavior entirely in output space. Instead of aligning diffusion latents or intermediate features, a lightweight pixel-space student learns the appearance and local structure of teacher outputs through RGB and spatial-gradient supervision, while paired high-resolution images and transcriptions provide reconstruction and recognition constraints. PixelSTR achieves over 100 parameter compression relative to the teacher restoration module. Under matched training, diffusion-derived supervision improves PSNR by 1.51 dB and OCR accuracy by at least 6.70 percentage points over supervised-only training, while replacing the diffusion targets with additional ground-truth targets does not reproduce these gains. These results suggest that generative restoration priors can be retained without retaining the generative model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.