Learning What to Align for Faster Diffusion Transformer Training
Abstract
Semantic emergence in Diffusion Transformers motivates alignment with external semantic representations to accelerate training convergence. However, observations suggest that such rigid alignment is often suboptimal, creating an inherent paradox: higher semantic accuracy does not necessarily lead to better generation fidelity. In this paper, we propose allowing the denoising objective itself to ascertain the specific facets of semantic information most beneficial to the generative process. Furthermore, we introduce a semantic internalization mechanism coupled with stage-wise training to estimate and reintegrate those task-relevant signals into the model’s intrinsic representation. Our acceleration regime prioritizes generation-critical features over rigid alignment with fixed-target semantic representations. Extensive experiments demonstrate that our approach consistently further accelerates training and improves generation quality across various model scales and resolutions, outperforming existing alignment-based baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.