PlanYourPixel: Structured Visual Planning for Fine-Grained T2I generation Control and Failure Localization
Abstract
Recent advancements in text-to-image (T2I) generation have improved visual quality, yet fine-grained control remains a fundamental challenge. Current layout-conditioned models struggle with numerical reasoning, attribute binding, and spatial grounding. These failures stem from the inherent ambiguity of natural language, where abstract prompts result in underspecified layouts. Even with LLM-based planning, the absence of an underlying scene logic causes models to assign coordinates before resolving the semantic relationships between objects, leading to floating items, incorrect spatial positioning, and attribute leakage. To bridge this gap, we introduce PlanYourPixel (PYP), an agent-based framework that utilizes specialized roles to decompose natural text into a structured scene representation. By resolving semantic constraints into a verifiable plan, PYP provides a unified foundation for two applications: layout-guided generation for Image synthesis and automated atomic question generation for fine-grained failure localization. We evaluated our paradigm on PYP-Bench, a diagnostic suite of 1,800 prompts and corresponding layouts. Our results demonstrated that PYP-guided image generation consistently outperforms standard layout baselines, while PYP generated 12,503 questions and formed a diagnostic framework that uncovers critical failure patterns that aggregate metrics typically conceal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.