TrajSP: Trajectory-Grounded Structured Prompting for Feedback-Guided Text-to-Image Generation
Abstract
While structured prompt enhancement enables image generation to benefit from visual knowledge and increased prompter capacity, existing distillation-based approaches require a separate training pipeline for an external LLM enhancer that lacks direct visual access at inference. We study whether unified multimodal models (UMMs) can eliminate the need for this separately trained enhancer by directly observing intermediate images and refining structured conditioning during denoising. Our framework, TrajSP, first applies a lightweight noisy understanding adaptation to enhance the understanding branch's capacity to interpret noisy intermediate images. During sampling, the adapted branch serves as a trajectory-grounded Polisher: at an early checkpoint, it processes the intermediate visual state jointly with the user prompt to predict a structured prompt comprising global scene fields, per-object elements with normalized bounding boxes, and inter-object relations. By grounding this representation in the evolving visual state, TrajSP provides targeted conditioning for subsequent denoising steps, with the aim of avoiding redundant prompt expansion. Joint understanding-generation finetuning optimizes shared parameters for both schema prediction from noisy states and schema-conditioned image generation. An optional late-stage Validator is designed to provide targeted feedback for subsequent refinement. Experimental results demonstrate improved semantic alignment and greater fidelity in object-level details, particularly attribute binding, without a standalone prompt enhancer. 15:52
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.