Mirrorrefine: Mirror Reflection Generation via Multi-agent Refinement with Spatial Priors
Abstract
Advancing visual generation toward world modeling requires more than photorealism: generated views must remain consistent with the underlying scene. In mirror-reflection synthesis, this requires inferring reflected content and spatial layout from the source scene, beyond the text–image alignment emphasized by general-purpose prompt-refinement methods. We propose MirrorRefine, a multiagent framework for mirror-reflection generation through multi-action evolutionary refinement. Starting from a generated or supplied image, the Reflection Optimization Controller analyzes the original request and visible content outside the mirror to infer intended reflected content and construct action-specific synthesis conditions, rather than treating potentially erroneous reflections as reliable targets. The Quality Evaluator provides a vision–language model with coarse reflectionlayout priors predicted by an offline-trained Reflection Spatial Geometry (RSG) module. Together with textual conditions, these priors guide the diagnosis of geometric, appearance, and illumination inconsistencies in the synthesized image. Based on this diagnosis, the evaluator selects among image generation, mirrorregion inpainting, and targeted editing, and returns structured revision suggestions to the controller. The controller updates the synthesis conditions and coordinates the selected operation, after which the resulting image is reassessed to guide subsequent refinement. All component parameters, including those of RSG, remain fixed during iterative refinement. Experiments on MirrorBenchV2, MSD, and CCMirror show complementary strengths: among the compared systems, the online configuration leads the reported geometry, lighting, and appearance scores, while the local configuration achieves the highest PSNR on MSD and CC-Mirror.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.