Bridging Scale-Equivariant Reconstruction and Photorealistic Generation for Blind Face Restoration
Abstract
In blind face restoration, a central challenge lies in reconciling two distinct objectives: scale-equivariant reconstruction and photorealistic generation. Reconstruction models naturally handle arbitrary degradation scales and preserve facial structure, but fail to synthesize realistic high-frequency details. Diffusion models provide powerful generative priors for detail synthesis, yet their fixed-resolution latent sampling process is not inherently compatible with continuous scale variations and often depends on expensive multi-step denoising. To unify the structural reliability of reconstruction with the perceptual realism of generation, we introduce FaceDNO, a novel paradigm that decouples the restoration process into two native mathematical spaces. Specifically, we formulate reconstruction in a continuous function space by employing a resolution-equivariant neural operator to project degraded inputs into a scale-robust structural anchor. We then formulate generation in the discrete latent space, integrating this continuous anchor into a diffusion prior via a manifold-preserving Latent Spatial Feature Transform (Latent SFT). To seamlessly bridge these two spaces, we formulate the operator's reconstruction in the latent space as an isomorphic linear structure to the diffusion forward process. By matching their marginal distributions via Signal-to-Noise Ratios (SNR), we analytically identify the optimal mid-timestep for generative refinement. Experiments demonstrate that FaceDNO effectively combines scale-equivariant reconstruction with photorealistic generation, achieving state-of-the-art fidelity and realism across diverse degradations with one-step diffusion refinement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.