SCOPE: Semantic-Bounded Credit Optimization for Prompt Engineering in Text-to-Image Generation
Abstract
Training a language model to rewrite prompts can improve text-to-image (T2I) generation while keeping the image generator frozen. In common reinforcement-learning formulations, a single image-level reward produces one trajectory-level advantage that is applied across the tokens of a rewritten prompt. This creates a credit-assignment gap: the units that users compose and that T2I systems must bind—objects, counts, attributes, and relations—are not the units that receive policy-gradient credit. We present SCOPE, a GRPO-style objective that uses semantic spans to structure this token-level policy update. For each rewritten prompt, SCOPE aligns generated spans with tokenizer positions, averages current-versus-old policy log-probability shifts within each span, and converts these shifts into detached weights normalized by valid-token count. The resulting weights modulate the token-level clipped policy surrogate, while the reward remains at the trajectory level. SCOPE therefore does not estimate the causal contribution of individual spans to the image reward; it provides a span-aware reweighting of policy updates. Across Qwen-Image, SD3, and FLUX.1-dev, SCOPE improves over vanilla GRPO on GenEval and T2I-CompBench, with mean relative gains of 9.48% and 7.62%, respectively. These results show that span-aware policy-update reweighting improves compositional alignment in the evaluated frozen-renderer settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.