TRACE: Trajectory-Routed Multi-Objective Alignment for Flow Matching Models
Abstract
Reinforcement learning from image-level rewards has become a standard approach for post-training flow-based text-to-image models. However, most existing methods optimize a single reward, and aligning multiple objectives such as compositional faithfulness, text rendering, and human preference remains challenging. The core difficulty is that multiple objectives conflict when optimized jointly on shared parameters. The difficulty appears along two temporal dimensions: within the denoising trajectory, objectives respond differently across phases; across optimization updates, simultaneous gradients can conflict and sequential updates can overwrite earlier gains. We therefore propose TRACE (Trajectory-Routed Alignment with Chunked Experts), which first estimates a per-phase multi-objective utility tensor via bucket-restricted probes and uses it to determine both where each objective is optimized and where each is protected. At the timestep level, TRACE converts this tensor into closed-form softmax routing weights that reweight the policy surrogate loss, while a timestep-conditioned LoRA provides phase-specific updates through a shared adapter. At the training-step level, TRACE trains focused single-objective chunks and protects idle objectives with profile-weighted anchor losses. Experiments on FLUX.1-dev and SD3.5-M with GenEval, OCR, and PickScore show simultaneous improvements over the base models and the highest reported point estimates among the evaluated multi-objective baselines on all three optimized metrics; CLIPScore and Aesthetic Score also remain above the corresponding base-model values. Code is available at https://anonymous.4open.science/r/TRACE-C427.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.