acceptodds
Under review as a conference paper at ICLR 2027

Mask-Aware Efficient Inference-Time Scaling for Image Inpainting

Abstract

Inference-time scaling has recently emerged as an effective paradigm for improving generation quality by allocating additional computation at test time. However, image inpainting remains largely unexplored in this setting, where complete-trajectory scaling is inefficient and structurally mismatched: it repeatedly scales the observed unmasked context and imposes a uniform computation budget on masked tokens with different refinement difficulties. Therefore, we identify two spatially non-uniform requirements for inpainting: inference-time scaling should exclude the observed unmasked region, while accounting for the varying difficulty across masked tokens. To this end, we propose a training-free, mask-aware framework for efficient inference-time scaling in image inpainting. For the unmasked context, Mask-Aware Initial Noise Construction and Anchor-Cached Background KV Cache remove observed-context redundancy by aligning background trajectories and reusing anchor background key/value states. For the masked region, Token-Aware Temporal Velocity Cache adapts masked-region updates by reusing velocity predictions for stabilized tokens and focusing computation on uncertain ones. Experiments on EditBench and BrushBench show that our method preserves scaling gains while reducing cost by up to and over standard scaling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.