RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Abstract
While image generation and editing have progressed significantly, existing models still struggle with ambiguous intent, complex reasoning, and out-of-distribution (OOD) tasks due to static knowledge gaps and reasoning limitations. To overcome these bottlenecks, we propose RS-Gen, a plug-and-play, training-free, multi-stage agentic framework for image generation and editing. By introducing a closed-loop "question-posing and problem-solving” mechanism, RS-Gen autonomously plans retrieval actions to bridge information gaps and execute deep logical reasoning. Experiments show that RS-Gen yields absolute gains of 0.31 (on a 0–1 scale) and 19.7 over Qwen-Image and Qwen-Image-Edit-2511 on WISE_Verified and RISEBench, respectively, achieving new SOTA results among open-source models and considerably pushing the capability boundaries of foundational backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.