SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
Abstract
The emergence of agentic image generation has expanded visual synthesis into a multi-stage process of knowledge grounding, reasoning, and iterative refinement. Yet, faithfully realizing complex user intent depends on preserving its underlying requirements across these stages. We refer to these requirements as semantic commitments and identify the Conceptual Rift as a break in their continuity: commitments may be resolved or checked locally without remaining explicitly tracked across grounding, generation, verification, and repair. We address this with SCOPE, a specification-guided skill orchestration framework that maintains semantic commitments in an evolving structured specification and conditionally invokes retrieval, reasoning, and repair skills to address unresolved or violated commitments. To evaluate whether generated images jointly satisfy these commitments, we introduce Gen-Arena, a human-annotated benchmark with entity- and constraint-level specifications, together with Entity-Gated Intent Pass Rate (EGIP), a strict entity-first pass criterion. Experiments on Gen-Arena, WISE-V, and MindBench show that SCOPE consistently outperforms the evaluated agentic baselines in overall performance across two generation backends, while reducing workflow orchestration cost and latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.