Finer Understanding, Finer Control: Towards Controllable and Generalizable Anomaly Generation
Abstract
Industrial anomaly generation has emerged as an effective approach to alleviate defect-data scarcity, yet existing methods, while generalizing to unseen object categories, remain limited in precisely controlling fine-grained anomaly attributes. In this work, we study zero-shot controllable anomaly generation, which requires both accurate attribute control and strong cross-category generalization. To avoid disturbing the pre-trained generalizable generative prior through direct generator training, our key insight is to cast controllable generation as a visual understanding problem with the generator kept frozen, and this can be inherently supported by leveraging a unified multimodal understanding–generation model (UMM). Based on it, we present the first Controllable and Generalizable Anomaly Generation framework, comprising grounded attribute understanding, on-policy generation alignment, and inference-time corrective steering. In particular, in the UMM, we first establish reliable anomaly attribute understanding and then use the grounded attribute space to align condition signal for the frozen generator via reinforcement learning. Our method successfully enables precise control over position, size, shape, and orientation on zero-shot object categories. Extensive cross-dataset experiments demonstrate state-of-the-art controllability and visual quality. We further show that stronger attribute understanding leads to more faithful generation and improves targeted augmentation for downstream anomaly detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.