InfinityFace: Unlocking Visual Autoregressive Text-to-Image Models for Blind Face Restoration
Abstract
Blind face restoration (BFR) is ill-posed, requiring a balance between perceptual realism and fidelity to the degraded input. We explore adapting Infinity, a bitwise visual autoregressive text-to-image model, to BFR. It is non-trivial as generation synthesizes plausible content from a generative prior, whereas restoration must retain the reliable input evidence. We propose InfinityFace, which reformulates open-ended next-scale generation as input-constrained next-scale restoration through two complementary forms of LQ conditioning. Specifically, a scale-dependent suffix controller combines the restored prefix with the corrected unresolved LQ features, preserving reliable LQ evidence while suppressing fine-scale feature mismatch. Meanwhile, scale-agnostic Text–LQ conditioning constructs an LQ visual memory and jointly injects it with text through multimodal cross-attention. Together, they anchor restoration to the observation while leveraging generative priors for realistic detail refinement. Experiments on synthetic and real-world BFR benchmarks show strong no-reference perceptual quality and competitive restoration performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.