RAEFace: Rethinking Latent Representations for Blind Face Restoration
Abstract
Most state-of-the-art blind face restoration (BFR) methods, whether codebook- or diffusion-based, rely on VAE-style latent representations. Representation autoencoders (RAEs), built on pretrained visual encoders, have shown strong performance in image generation but remain largely unexplored in restoration-oriented low-level vision. We show that RAE latents better preserve facial semantics under severe degradation and exhibit stronger semantic compatibility between high-quality (HQ) and low-quality (LQ) observations than VAE latents. In light of this, we propose RAEFace, a framework that unifies HQ target modeling and LQ condition injection in a shared RAE space to produce semantically faithful base restorations. To seek perceptually stronger yet faithful solutions, we introduce Perceptual Improvement via Extrapolation (PiE), a training-free inference strategy that performs stepwise extrapolation during flow-matching sampling. Using the base prediction as a fidelity anchor, PiE extrapolates beyond it along the direction from the degraded latent to its restored counterpart. Our analysis links the effectiveness of this direction to HQ–LQ semantic compatibility, which helps preserve facial content during extrapolation. Experiments on real-world and synthetic benchmarks demonstrate that RAEFace achieves state-of-the-art perceptual quality with competitive fidelity. Results on inpainting and colorization further suggest the broader potential of RAE representations for related low-level vision tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.