acceptodds
Under review as a conference paper at ICLR 2027

DIFFSCHED: Closed-Loop Semantic Computation Scheduling for Efficient Compositional Text-to-Image Generation

Abstract

Text-to-image diffusion models still struggle with prompts that combine multiple objects, attributes, counts, and relations. Regional control can improve compositional accuracy, but fixed schedules may waste computation on resolved requirements while providing insufficient intervention for difficult ones. We propose , a closed-loop scheduling framework that adapts regional computation and semantic review during generation. Before sampling, role-specialized agents decompose the prompt into objects, relations, spatial regions, and an initial generation plan. During sampling, a lightweight trigger tracks the magnitude and recent trend of disagreement between regional and global denoising predictions. At eligible checkpoints, a low trigger score stops the current regional interventions without invoking semantic review; otherwise, a Review Agent diagnoses unresolved requirements and the Scheduler updates the next intervention interval. The residual guides execution rather than certifying semantic correctness, while high-risk requirements retain mandatory checks. On GenEval 2 and T2I-CompBench++, achieves a better quality–computation trade-off than fixed scheduling heuristics, using approximately 42–56% fewer denoiser evaluations while improving compositional scores. These results show that selectively allocating regional computation and semantic feedback can improve compositional generation without uniformly increasing inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.