acceptodds
Under review as a conference paper at ICLR 2027

Debate-Driven Generation: Multimodal Agent Debate for Iterative Text-to-Image Generation

Abstract

Recent advances in text-to-image generation have placed growing emphasis on reasoning for better instruction following and refinement. Nevertheless, the need to transform sparse language into dense visual content makes contextual disambiguation and elaboration a central bottleneck beyond cross-modal alignment. Therefore, we introduce the Multimodal Agent Debate (MAD) framework that adapts multi-agent debate to iterative text-to-image refinement by addressing three gaps, namely evolving debate topics, the mismatch between semantic precision and execution feasibility, and the need for factual rather than reflective moderation. MAD maintains (i) an explicit roadmap of verifiable criteria grounded in the current image, (ii) uses a Regulator to decide whether debated instructions are executable by a downstream editor or require re-debate with pragmatic constraints, and (iii) employs a cross-modal Moderator to verify prompt satisfaction and produce verifiable success conditions. Experiments on GenEval, T2I-CompBench, and TIIF-Bench show consistent overall gains in instruction following across diverse image generators. In addition, MAD also improves instruction executability and reduces repetitive refinements, supporting MAD as a training-free refinement method for T2I. The source code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.