acceptodds
Under review as a conference paper at ICLR 2027

Build the Scaffold, Then Generate: Explicit Spatial Planning for Spatially Faithful Image Generation

Abstract

Text-to-image models produce high-fidelity images but remain unreliable in realizing complex spatial compositions. This limitation is more pronounced in open-source models, with a notable gap to proprietary models. We propose ScaffoldGen, an agent-driven framework that separates spatial planning from visual synthesis. Given a prompt, the planning agent parses constraints from the prompt and constructs a global layout plan in HTML, which is rendered into a visual scaffold. The generator is guided by the visual scaffold to synthesize detailed visual appearance. For systematic evaluation, we introduce SLAB, a Spatial LAyout Benchmark covering ten task categories involving complex spatial composition and layout organization. The benchmark assesses semantic and spatial constraint satisfaction through dependency-aware question-answer accuracy (QAA). On SLAB, ScaffoldGen outperforms all evaluated open-source methods and achieves competitive performance with certain proprietary models. Further evaluation on SpatialGenEval also shows consistent improvements over the baseline, confirming that the enhanced spatial awareness generalizes beyond SLAB. Together, the results demonstrate the effectiveness of ScaffoldGen and point to a broader research direction: leveraging the spatial reasoning capabilities of LLM agents to compensate for the limited spatial perception of generation models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.