CtrlGen: Optimal Control for Text-to-Image Generation
Abstract
Flow-based text-to-image models achieve strong generation quality and efficiency, yet conditional generation still relies heavily on classifier-free guidance. Existing guidance methods mainly adjust when or how strongly the conditional correction is applied, while its direction and magnitude are usually regulated together. This coupling limits local strength control, and unconstrained spatial and temporal magnitude variations can produce excessive responses during sampling. We identify guidance magnitude as the key variable for controlling strength while preserving the model-predicted direction and suppressing excessive local magnitude variations. Based on this observation, we propose CtrlGen, a training-free framework that decomposes the conditional correction into local direction and magnitude and treats the latter as a state-dependent control variable. At each sampling state, CtrlGen solves a regularized convex control problem that balances guidance fidelity, control energy, spatial smoothness, and state consistency. The resulting optimality condition is a screened elliptic equation with an explicit spectral solution, enabling frequency-selective magnitude control without iterative optimization. The optimized magnitude is recombined with the preserved direction to form a feedback guidance rule without modifying the pretrained model. Extensive experiments show consistent improvements over existing guidance methods in generation quality, text alignment, and preference-oriented metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.