IN-PLACE DEFECT EDITING VIA TIMESTEP-CONSISTENT MASK DIFFUSION
Abstract
Few‑shot defect generation mitigates the scarcity of labeled anomalies in industrial visual inspection by synthesizing realistic and diverse defects from a few real examples. Two paradigms emerge: good‑to‑defect diffusion renders a defect onto a clean template, carrying correspondence risk and domain gap; defect‑to‑defect editing preserves context but rearranges observed pixels. We find that neither fully preserves defect realism, and we recast augmentation as in‑place defect‑to‑defect editing, re‑synthesizing the defect at its original location under diffusion‑governed fidelity rather than template correspondence. Mask‑guided diffusion, however, trains a soft region weighting but samples a hard binary cut, so the imposed boundary is one the model never learned to respect. We couple the mask to the diffusion timestep: a Timestep‑adaptive Dynamic Mask Region (TDMR) loss softens the constraint at high noise and sharpens it at low noise, while stepwise blended sampling re‑anchors the background at every step with the same mask. Treating the mask as an evolving schedule shared between training and sampling, we identify a critical yet under‑explored principle: training‑sampling mask consistency, whose violation yields boundary artifacts even when the mask is exact, and whose hard boundary is the limit of the softened training schedule. On MVTec AD, a detector trained solely on our synthetic data attains the best localization accuracy among protocol‑matched augmentation baselines, consistent with this principle; mixing with real data yields further gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.