UGenIR: Understanding-Conditioned Discrete Generation for Open-World Image Restoration
Abstract
All-in-one image restoration is commonly formulated as a closed-set multi-task regression problem: a single model is trained to map a degraded input to a restored output in one shot. Such regressors often fail under open-world corruptions, where degradation types, severities, and especially their compositions shift beyond the training distribution, making pixel evidence ambiguous and leaving little room for revision. We propose UGenIR, which reframes restoration as understanding- conditioned constrained generation. Given a degraded image, UGenIR first infers an input-consistent semantic description that acts as a language- level prior to constrain global content. Restoration is then performed in discrete token space via MaskGIT-style parallel token completion, where low-confidence tokens are progressively re-masked and refined for effi- cient iterative recovery of fine details. To further improve robustness under distribution shift, we introduce a critic-guided self-correction loop: a lightweight critic predicts spatial error maps during training, and at test time its feedback is converted into token masks to trigger targeted rollback and regeneration on likely-error regions. UGenIR achieves con- sistent gains on multi-task restoration and challenging OOD degradation compositions, improving both perceptual realism and semantic/content consistency with efficient discrete refinement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.