Plan Before You Denoise: Semantic Availability in Image Generation
Abstract
A text-to-image model can render a visual answer far more reliably than it can execute the instruction that determines it. On controlled Boolean-program prompts (SYNBOOL), FLUX renders the correct colour on only 27.1% of unresolved instructions at 15 operations (collapsing from 60.8% at 3 ops to 24.8–27.1% for ops), but reaches 80.7% when the answer is supplied. Across program depths, sampling budgets, and two model families, unresolved generation collapses to chance levels on multi-step programs, dominated by superficial prompt-position heuristics. We formalize this execution gap through three complementary computational constraints: (1) Content: grounded in usable information (Xu et al., 2020), an exact decoder-relative velocity-risk law shows that when semantic readouts cannot resolve the instruction, the population-optimal restricted velocity collapses to uninformative drift, forfeiting conditional guidance despite statistically complete input. (2) Adaptive compute: a stylized random-oracle benchmark isolates serial depth from parallel width, establishing that with fewer than adaptive rounds, achieving nontrivial advantage requires total queries. (3) Timing: a finite-Euler controllability bound proves that late disclosure monotonically contracts the remaining Gramian, escalating the theoretical minimum steering energy, while empirical evaluations show that standard linear guidance incurs severe control inefficiency when delayed. Upfront planning resolves this bottleneck by paying the serial span in autoregressive reasoning tokens and caching the visual decision from the very first update. These constraints provide insight into complementary failure modes for one-shot planning (single-pass degradation under spatial load) and iterative painting (compounding routing errors across sequential hops). Combining them into a constraint-guided hybrid PLAN+PAINT strategy pays serial span in an upstream planner and spatial work across incremental rendering passes, boosting exact placement on dense relational grids to 0.623–0.773 (vs. 0.401–0.525 for Plan and 0.159–0.529 for Paint), reaching 0.745 on GenEval spatial prompts (vs. 0.395 unresolved), and achieving 0.770 on T2I-ReasonBench entity prompts (vs. 0.345 unresolved, matching an answer-supplied oracle of 0.775).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.