RISA: Recursive Improvement for Scene-Consistent Visual Anomaly Synthesis
Abstract
Visual anomaly synthesis provides a practical way to supplement scarce real-world anomaly data for industrial inspection and safety monitoring. Although instruction-guided image editing enables controllable anomaly generation from natural-language descriptions, existing methods typically rely on one-shot generation or loosely constrained iterative refinement, making it difficult to correct synthesis errors without disturbing the spatial structure and unaffected content of the original scene. We introduce RISA, a multimodal agent framework for Recursive Improvement for Scene-Consistent Anomaly Synthesis. RISA formulates anomaly synthesis as a progressive process of observation, diagnosis, and repair. At each iteration, the agent evaluates the synthesized image at both global and local scales, identifies discrepancies in anomaly appearance, scale, placement, and integration with the surrounding scene, and adaptively revises its editing strategy based on feedback from previous attempts. To maintain scene consistency throughout refinement, RISA combines multimodal scene understanding with spatially constrained image manipulation, including cropping, localized editing, and compositing, enabling the agent to jointly determine what anomaly to generate, where to place it, and how to integrate it into the original scene while minimizing unintended changes. This recursive inference-time improvement process requires no parameter updates to the underlying models. Experiments across industrial inspection and safety monitoring scenarios demonstrate that RISA consistently improves anomaly fidelity while better preserving the structural and visual consistency of the original scene.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.