UAI-Net: Uncertainty-Aware Image Inpainting with Learned Adaptive Refinement Depth
Abstract
Most inpainting models fill in missing regions without indicating which parts of the fill are likely to be wrong, and those that do estimate a confidence map use it only to order pixels within a fixed number of refinement passes. This paper makes the model estimate its own difficulty and act on it. UAI-Net predicts a coarse spatial uncertainty map, trained to reproduce the rank ordering of the model's reconstruction error, and uses it twice: to weight the fusion of local, global-attention and boundary features, and to decide, through a straight-through halting head, how many refinement steps (one to four) each location of the refinement grid receives. Training the halting head is the main difficulty: with a continuous depth target its logits collapse to zero and the executed depth is decided by floating-point rounding; an integer target with a small confidence penalty removes the problem. On CelebA-HQ, Places2 and FFHQ (256x256, free-form masks over 10-60% of the image, 1,500 test images each) the map correlates with the true per-pixel error at a within-image Spearman rho of 0.60-0.67, and the model's predicted variance detects badly filled pixels better than the flip-disagreement or snapshot-ensemble uncertainty of stronger inpainters, in one forward pass. On CelebA-HQ the least-uncertain tercile receives 1.7 refinement steps and the most-uncertain 3.2 (2.54 on average), matching the quality of uniform four-step refinement; the block is still evaluated densely, so wall-clock time is unchanged. Against a PSNR-optimised variant at equal PSNR, UAI-Net gives 8-15% lower LPIPS and 10-17% lower FID.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.