acceptodds
Under review as a conference paper at ICLR 2027

T-MOG: Temporally Structured Multi-Objective Guidance for Text-to-Image Diffusion Models

Abstract

High-quality text-to-image generation must satisfy multiple objectives, including semantic consistency and aesthetic preference. However, existing multi-objective guidance mainly assigns weights across objectives by introducing multiple rewards at each step, implicitly assuming all objectives remain useful throughout generation. This parallel modeling overlooks the coarse-to-fine dynamics of diffusion, where semantics and spatial structure are usually established early, while local details are refined later. Thus, it remains underexplored whether different objectives should continuously participate in optimization. Continuous joint optimization may cause conflicts within the same stage and allow later detail refinement to disrupt earlier semantic structures. To address this, we propose T-MOG(Temporally Structured Multi-Objective Guidance), a training-free multi-objective guidance framework with temporal decoupling. Unlike existing methods that focus on coordinating objectives simultaneously, T-MOG redefines multi-objective guidance from an orthogonal perspective. It adaptively determines when and in what order different objectives should be optimized according to diffusion dynamics. Specifically, it dynamically selects the most suitable guidance objective for each generation stage, enabling each objective to act when most effective. We further introduce a cross-stage memory protection mechanism that monitors the influence of a new objective’s guidance direction on previous objectives after switching. It suppresses updates that may degrade earlier gains, thereby balancing new-objective improvement with the preservation of prior optimization. Comprehensive experiments demonstrate that T-MOG consistently delivers more balanced joint improvements than single-objective guidance and static reward fusion while preserving previously optimized objectives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.